Meta's 9B user simulator beats Claude-Opus-5 on four benchmarks and lifts agent RL

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng, Wancen Mu, Yue Cao, Shengjie Bi, Yun He, Changdae Oh, Deren Lei

cs.AI

2026-10-07

Meta trains MIMESIS, a 9B user simulator that beats Claude-Opus-5 on four benchmarks; agents trained against it beat GPT-5.5-trained agents under all nine unseen evaluation users.

What problem this solves

Training interactive agents needs someone on the other end of the conversation. Human trajectories are expensive and hard to scale, so the standard shortcut is prompting an off-the-shelf assistant LLM to role-play the user. The paper quantifies how far off that is: a fixed GPT-5.5 agent solves 63.6% of τ-bench tasks with real users; three frontier API simulators push the rate to between 82.4% and 84.4%; pretrained Qwen3.5 models drag it down by 14.9 to 16.0 points. Assistants are too cooperative and volunteer too much information, which makes tasks easier. Raw pretrained models make them harder.

The simulator decides the distribution of trajectories an agent sees during evaluation and RL, so difficulty distortion poisons the reward signal directly: policies trained in an environment nearly 20 points easier than reality will break on real users. Purpose-built simulators stay much closer to the human reference, deviating by at most 10.3 points, versus 1.8 and 2.4 for the two MIMESIS sizes. The paper's question is whether one model can be both behaviorally realistic and a good training environment, and it answers with a two-stage pipeline.

Method

MIMESIS comes in 4B and 9B sizes on Qwen3.5 backbones. Stage one trains the user in three steps:

The behavior task is the distinctive part. From ThoughtTrace the authors derive 13 recurring user behaviors; the most frequent in the annotated sample are hidden evaluation criteria (40.7%), incremental goalpost shifting (13.4%), and refusal to cooperate with clarifications (11.1%), alongside fabricated premises, self-contradiction, and entitled pressure. Each rollout is grounded in an ABCD customer-support scenario with structured records and policies: withheld information must actually exist in the record, and an infeasible request must actually conflict with policy, so the simulator cannot just hallucinate difficulty. The reward weights behavior exhibition, timing, naturalness, and scenario consistency at 2:1.5:1:0.5, and is halved when the model describes the behavior instead of acting it out. A third of the data carries no imposed behavior.

Stage two freezes the simulator and trains an agent with multi-turn GRPO against it, plus a new objective: Coached On-Policy Self-Distillation (CSD). After each sampled agent reply, a coach model reads the public history, the reply, the simulator's private thought, and the user's next utterance, then writes a short note on how the reply could have served the user better. A stop-gradient teacher copy of the policy, conditioned on the note, scores the same sampled reply; the per-token log-probability gap between teacher and student passes through a sigmoid into weights on a likelihood loss added to GRPO. Unlike SDPO, CSD re-scores the already-sampled reply instead of decoding a refined replacement. At deployment the agent conditions only on the public dialogue.

Results

Simulator side, four benchmarks:

MetricMIMESIS-9BBest baselineGap
SOUL-Index65.7Claude-Opus-5 64.9+0.8
RealUserSim PT3 fidelity94.0Claude-Opus-5 80.6+13.4
τ-USI (five-component)80.17Osim-8B 80.44−0.27
SimulatorArena Turing distance38.7Claude-Opus-5 42.3−3.6

The 4B model scores 63.7 on SOUL, also above every released simulator, and the two sizes gain 15.2 and 18.0 points over their own backbones. On RealUserSim the margin comes from interaction and information flow (91.3 vs 69.5) and pacing (91.2 vs 72.5); persona agreement is near ceiling for several models, so the lead reflects how information unfolds across turns. On τ-bench, MIMESIS-9B posts the lowest expected calibration error, 0.069 versus 0.135 for Osim-8B and 0.188 for GPT-5.5, meaning it best reproduces the task difficulty real users impose.

Agent side: across eight environments (three held out) and nine evaluation user models, none seen in training, GRPO with a GPT-5.5 user averages 26.10. Swapping in MIMESIS-9B raises the mean to 29.54, with gains of 1.46 to 4.55 points, positive under all nine users. Adding CSD reaches 31.09, again improving under all nine. The last two conditions share the same simulator, so the 1.55-point increment is attributable to the coaching objective alone.

Why it matters

For teams building agent RL pipelines, this turns simulator choice from a style question into a correctness question: assistant models playing users systematically inflate success rates, corrupting both leaderboards and reward signals. A 9B open-weight model that beats frontier APIs at this job lets the user side of the gym run locally and scale without per-rollout API cost.

The CSD recipe travels beyond this paper. It shows how to convert training-only privileged information (here, the simulated user's reasoning) into dense token-level supervision while the deployed policy depends on nothing privileged, which applies anywhere the training environment knows more than the agent will at test time.

The margins on simulator benchmarks are mostly single digits outside RealUserSim. That is evidence specialization pays, not that the model is generally stronger than frontier APIs.

Limitations

The authors' own accounting: a Turing distance of 38.7 is still far from zero, where the judge could not tell simulated from human at all; frontier models keep the lead on SOUL's social-simulation and role-play axes and on SimulatorArena's explicit style ratings; and Osim-8B leads the communication-style dimension of τ-USI (62.12 vs 51.75). The τ-USI variant used here also drops the survey component of the original six-part index because the annotations were unavailable.

CSD's feedback source carries a circularity risk the paper mentions only in passing. Coaching notes rest on the simulator's own generated thoughts, not on any human's internal state; if the simulator misreads user intent, CSD distills the wrong adaptation into the agent at token-level density. The generalization evidence also comes entirely from nine other simulators, with no direct comparison against live human users; alignment with humans is argued only through the ECE proxy. The 13-behavior taxonomy derives from GPT-5.6 annotation of 2,155 conversations, and the paper itself notes the frequencies describe the annotated sample, not population prevalence. Reproduction is not cheap either: mid-training took about 4.8 days on 32 A100s, and RL another 500 steps on 16.

Terms

Source

Related papers

All paper explainers