Question's Gambit tops BrowseComp-Plus recall with 96.6% using only BM25
CShorten30 · x · 2026-09-25
The Question's Gambit module for search agents tops the BrowseComp-Plus recall leaderboard using only BM25, achieving 96.6% recall with 6x fewer agent search calls. It improves the agent's first retrieval move: decomposing the question into clues, reformulating them into complementary searches, and reranking before the iterative loop. With gpt-5.5, answer accuracy rises from 83.1% to 90.5% over the strongest agentic baseline Pi-Serini; gains also transfer to MultiHop-RAG. Paper: arXiv:2609.14412.
More from coding & agent
- SkillRL (NeurIPS 2026): 7B model beats GPT-4o by 41% via recursive skill evolution — cihangxie · 2026-09-25
- Weave Code Max offers $50+ of coding model usage for $10/month — ycombinator · 2026-09-25
- Mad science: Jev autopilot lands a plane in a terminal flight simulator — bilawalsidhu · 2026-09-25
- Greptile reviewed 395K PRs at NVIDIA, cutting merge time from 24h to 6h — ycombinator · 2026-09-25
- Open-source Claude skill turns one prompt into stop-motion claymation films in Blender — angrypenguinPNG · 2026-09-25
- Descript launches MCP to edit video from inside Claude, ChatGPT, and Codex — descript · 2026-09-25