HAI.AI research · study 01 · snapshot 2026-07-24

MediationBench

A synthetic-conflict testbed for building and evaluating AI mediators. We replay the same authored fixture and starting transcript, vary the mediator's policy and information, and measure what changes.

The measurement problem

A fluent intervention is not the same as an effective mediator.

Mediation unfolds through a conversation. The disputants, the information available before the room, the simplest active control, the stopping rule, and the evaluator can all change the apparent result.

MediationBench makes those choices visible. Its synthetic fixtures provide private interests and matched starting points, so policies can be compared across otherwise matched simulated runs.

What the current evidence changes

Three controls against an easy story.

01

The simulated disputant is part of the instrument.

One cooperative configuration compressed skilled mediators into a narrow band. A high-resistance configuration made policy differences measurable. That is configuration sensitivity, not proof that one model is "real."

02

Reflection is an active control, not a placebo.

Acknowledgment and reflective listening can change a dispute. Skill is therefore measured above a model-matched reflective policy, not above silence alone.

03

A failed generalization belongs in the result.

In the 2026-07-24 analysis, the replicated Qwen core and conditional Kimi breadth evidence retained a public-context skill increment; gpt-oss did not show the same increment. These model-specific results do not establish a backbone-universal effect.

Full bilateral brief

Private scenario facts supplied.

Both parties' researcher-authored private facts are provided verbatim alongside public context.

Public-context only

Relevant constraints must surface.

The mediator receives the public scenario, shared facts, and dialogue, but no private-fact packet.

Information is a treatment

Preparation is not a nuisance variable.

The bilateral brief is a controlled analogue of one advantage that intake, interview, or caucus may provide; it is not an observed human preparation process. The study asks what that information changes and whether a skilled policy can elicit absent constraints.

See the prospectively frozen 2 × 2 study

For mediator and foundation-model builders

Use synthetic conflict as a development loop.

Test whether a prompt adds value beyond reflection. Compare full-brief and public-context policies. Inspect failures by conflict type. Re-score fixed transcripts without regenerating them. Generate controlled trajectories for evaluator and mediator development.

Claim boundary

Models in conflict, not people in treatment.

Every party in this study is simulated. MediationBench measures controlled model behavior; it does not establish safety, legal validity, or effectiveness with human participants. Human calibration and same-backbone participant manipulations remain required follow-up work.

Human Assisted Intelligence is a Public Benefit Corporation

Build AI that can help without taking the human's place, and measure it under conditions that make failure visible.

The product lives at hai.ai. This site publishes MediationBench, its philosophy, and its bounded evidence.