Question's Gambit tops BrowseComp-Plus recall with 96.6% using only BM25

CShorten30 · x · 2026-09-25

The Question's Gambit module for search agents tops the BrowseComp-Plus recall leaderboard using only BM25, achieving 96.6% recall with 6x fewer agent search calls. It improves the agent's first retrieval move: decomposing the question into clues, reformulating them into complementary searches, and reranking before the iterative loop. With gpt-5.5, answer accuracy rises from 83.1% to 90.5% over the strongest agentic baseline Pi-Serini; gains also transfer to MultiHop-RAG. Paper: arXiv:2609.14412.

Original post →

More from coding & agent

coding & agent channel →