Agents judge tool results useless 97-100% of the time yet rarely stop: NTU study

nanyang-technological-university-singapore · hf · 2026-10-07

An NTU paper documents a counterintuitive gap: agents judge useless tool results but don't act on it. In a retrieval environment with controlled source failures, the seven agents tested call a failing source's results useless 97-100% of the time, yet most rarely stop on that judgment. Prompt cues change when they stop, not what they stop on; permission to answer from memory and a reasoning mode trigger early stops regardless of evidence; a stated budget just pushes 7-8B models' stops to the deadline.

The only fix that works: harness-enforced integration—forcing the agent to answer after five consecutive results it judged useless. This raises failing-source success for every model and keeps the stopping point fixed when budget doubles. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.

Original post →

More from coding & agent

coding & agent channel →