Amazon paper: LLM judges rate 57.5% of failed agent tasks as satisfactory, flip close rankings 31% of the time
dair_ai · x · 2026-09-14
A new Amazon paper stress-tests the common agent-eval setup of an LLM user simulator plus an LLM judge scoring transcripts, and finds it fails in two specific ways: 1) Satisfaction doesn't track success — 57.5% of conversations raters marked satisfied had failed the customer's task. 2) Close calls go wrong — rankings hold across agents of very different ability, but among near-equal agents the judge picks the lower-reward one on 31% of pairs (vs under 1% for far-apart pairs). The study covers 25 agents from six providers on tau2-bench and SimulatorArena; judges also favored agents from their own model family. The fix is cheap: a judge-free completion bit catches truncation regressions, and judges should be trusted only after calibration against ground truth.
More from coding & agent
- AI Filmmaker Spent $4.5k and 847 Takes to Make a 22-Minute Short with Seedance — DavidmComfort · 2026-09-14
- Evomap launches EvoX beta: one swarm agent that codes, researches, and builds slides — CodeByPoonam · 2026-09-14
- Dev Proposes AI-Safety Contractor Board With Cross-Model Verified Receipts — EricBuess · 2026-09-14
- LangChain ships Managed Deep Agents, evolving from LangSmith Deployment into a full agent runtime — hwchase17 · 2026-09-14
- New Podcast with Cloudflare Engineer: AI Coding Agents, Local Models and Agent Harnesses in Practice — Arindam_1729 · 2026-09-14
- Dev on AI coding: explore with ambiguity, but know exactly what to build before building — _Stocko_ · 2026-09-14