Amazon paper: LLM judges rate 57.5% of failed agent tasks as satisfactory, flip close rankings 31% of the time

dair_ai · x · 2026-09-14

A new Amazon paper stress-tests the common agent-eval setup of an LLM user simulator plus an LLM judge scoring transcripts, and finds it fails in two specific ways: 1) Satisfaction doesn't track success — 57.5% of conversations raters marked satisfied had failed the customer's task. 2) Close calls go wrong — rankings hold across agents of very different ability, but among near-equal agents the judge picks the lower-reward one on 31% of pairs (vs under 1% for far-apart pairs). The study covers 25 agents from six providers on tau2-bench and SimulatorArena; judges also favored agents from their own model family. The fix is cheap: a judge-free completion bit catches truncation regressions, and judges should be trusted only after calibration against ground truth.

Original post →

More from coding & agent

coding & agent channel →