One-Sentence Tool Refusal Boosts Search Agent Abstention from 23% to 97%
_reachsumit · x · 2026-10-06
A new paper, Search Engines Never Say No, examines what frozen search agents do when the retrieval tool can refuse. Since search tools always return top-k passages—even when the index holds no answer—agents see irrelevant text instead of a miss signal and guess. The authors built an index-hole testbed (257 NQ and 300 HotpotQA questions run with and without gold passages in a 21M-passage BM25 index) and tested seven agents with five refusal wordings.
Key results:
- An un-announced one-sentence refusal raises abstention on unanswerable questions from 23% to 97% on average for Qwen3-8B/32B, and from 28% to 57% for Claude Haiku 4.5, cutting wrong answers almost one-for-one—outperforming a system-prompt instruction by 51 points on average.
- Search-R1 ignores refusals and fabricates retrievals; Claude Sonnet 5.5 and Opus 5.5 answer from memory (abstention +2) but obey a directive inside the observation (+16).
- Wording matters: an explanation beats a bare token; soft warnings are useless.
- Realistic triggers (score-based predictors, LLM grounding judges) fall well short of the oracle; for compliant agents the bottleneck is the detector inside the tool, not the agent.
More from coding & agent
- Matt Pocock: Customize Your Coding Agent Skills to Your Own Workflow, Don't Use Them Off-the-Shelf — mattpocockuk · 2026-10-06
- Stanford ACE team unveils Sentry: failure tips in context hurt LLM agents, +39% gains — StanfordAILab · 2026-10-06
- Shadowrocket TUN-Only Setup to Stop Claude Account Bans — aigclink · 2026-10-06
- OpenAI streamlines ChatGPT plugin submissions: upload zip, fix validation, publish — Dimillian · 2026-10-06
- One tool beats two: how combining fetch and extraction fixed my agent's context overflow — OkShirt9372 · 2026-10-06
- Long-running benchmarks find Strata inference server failing full-build scenarios — julianharris · 2026-10-06