Decagon on GEPA-GAN: Simulated Users That Are Too Cooperative Are Skewing Agent Evals
kastnerkyle · x · 2026-09-13
Decagon shares an article, "GEPA-GAN: Teaching AI to Sound Human," noting that customer-support agents are evaluated against simulated users, making the simulator part of the benchmark. If the simulated customer is cleaner, more patient, or more cooperative than real users, evals become systematically optimistic; the piece explores making simulated users sound human to keep benchmarks valid.
More from coding & agent
- Full SlopCodeBench run: GPT-6 Astra edges GPT-5.6 Sol and GLM 5.3, but not by much — cedric_chee · 2026-09-13
- After the OpenAI-HF fallout, someone built collusion.gg — a forum only AI agents can use — aflyr3 · 2026-09-13
- 402cron: an MCP server for scheduled, HMAC-signed HTTP delivery to agents — pay402 · 2026-09-13
- Tiling window managers shine on 21:9 ultrawide: browser beside Codex or Cursor — mark_k · 2026-09-13
- Could AI agents peer-review each other? Reddit debates an 'Agent Passport' idea — Most-Bat2916 · 2026-09-13
- Claude Code system prompt tells Opus 5 to avoid excessive self-correction, but not Opus 4.8 — repligate · 2026-09-13