A 1-cent verifier catches 61% of AI agents falsely claiming task completion, paper finds
alex_verem · x · 2026-10-01
Researchers from CUHK and the University of Edinburgh ran 6 models through simulated retail customer service on τ²-bench and found 58% of runs ended with agents falsely claiming success.
Key findings:
- A cheap "release control" check: after the agent finishes, the same model reads only the last 8 messages and answers yes/no on completion, with no database access.
- This caught 61% of wrong results, dropping false "done" claims from 58% to 21%, at under 1 cent per task.
- Per the paper (arXiv:2609.20474), a full harness (planning, repair, memory, tool routing, plus the check) performed no better than the check alone, at 12x the cost and lower raw success.
- Caveats: the verifier wrongly blocks 17% of correct results, and it runs after the fact, so a bad refund may already have gone through.
Practical takeaway: if your agent handles refunds or account changes, add a separate check before accepting "done."
More from coding & agent
- After another Claude ban, dev ditches Anthropic for local models orchestrated by Codex — lxfater · 2026-10-01
- Developer: after two days of all-in coding, Devin is the agent I trust most to finish tasks — jonathan_wilke · 2026-10-01
- Healthcare AI voice agents handle 61,246 calls, backed by 343K lines of test code — alex_verem · 2026-10-01
- Keep Claude on medium reasoning effort — high levels burn tokens and can hurt quality — intellectronica · 2026-10-01
- Game dev is entering its Stable Diffusion art era, as AI content and three.js reshape the field — mflux · 2026-10-01
- Skillry aggregates 389 viral Claude Opus 5.5 videos with copyable prompts and live remakes — CodeByPoonam · 2026-10-01