Decagon on GEPA-GAN: Simulated Users That Are Too Cooperative Are Skewing Agent Evals

kastnerkyle · x · 2026-09-13

Decagon shares an article, "GEPA-GAN: Teaching AI to Sound Human," noting that customer-support agents are evaluated against simulated users, making the simulator part of the benchmark. If the simulated customer is cleaner, more patient, or more cooperative than real users, evals become systematically optimistic; the piece explores making simulated users sound human to keep benchmarks valid.

Original post →

More from coding & agent

coding & agent channel →