Agents judge tool results useless 97-100% of the time yet rarely stop: NTU study
nanyang-technological-university-singapore · hf · 2026-10-07
An NTU paper documents a counterintuitive gap: agents judge useless tool results but don't act on it. In a retrieval environment with controlled source failures, the seven agents tested call a failing source's results useless 97-100% of the time, yet most rarely stop on that judgment. Prompt cues change when they stop, not what they stop on; permission to answer from memory and a reasoning mode trigger early stops regardless of evidence; a stated budget just pushes 7-8B models' stops to the deadline.
The only fix that works: harness-enforced integration—forcing the agent to answer after five consecutive results it judged useless. This raises failing-source success for every model and keeps the stopping point fixed when budget doubles. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.
More from coding & agent
- Agentic AutoRAG: LLM Agents Diagnose Retrieval vs Generation Failures to Tune RAG Pipelines — _reachsumit · 2026-10-07
- Cursor adds Cloud Agents API endpoints for environment builds with status and error codes — tetsuoai · 2026-10-07
- Figure CEO: filling Vietnam's brutal visa form was our AGI test — now an agent passed it — adcock_brett · 2026-10-07
- Vite+ 1.1 released: 24% faster vp dev startup, 30% less memory, clearer prompts — irvinebroque · 2026-10-07
- solid-yield Brings Generator-Based Type-Safe Components to Solid 2 as AI Sparks a Yield Renaissance — samgoodwin89 · 2026-10-07
- Rex, a Coding Agent Multiplexer, Opens Mailing List Invites for Early Testing — DanielLockyer · 2026-10-07