Paper: Sol 5.6 Achieves 46% Accuracy on BrowseComp-Plus Without Search
beirmug · x · 2026-08-28
A paper by Sahel Sharify et al., titled "Projecting BrowseComp-Plus onto ClimbMix," examines the realism of agentic search benchmarks.
Background:
- The original BrowseComp-Plus benchmark relies on a fixed corpus of 100K documents curated for specific queries, lacking the noise of real-world data.
- The authors project these questions onto NVIDIA ClimbMix (400B tokens, 553M documents), a real-world general corpus, using a pipeline that decomposes questions into atomic reasoning hops and verifies evidence.
Key Findings:
- Data indicates that the Sol 5.6 model achieves a 46.1% native accuracy on BrowseComp-Plus using parametric knowledge alone (without search tools).
- After projection onto ClimbMix, retrieval difficulty increases significantly, causing the strongest agent's score to drop by 5 points, better reflecting real-world challenges.
More from Research
- Agent Benchmark Includes Real-World Noise: 300 Phishing Emails and 10 Error Classes — testingcatalog · 2026-08-29
- NeurIPS Sim2Science workshop extends submission deadline to Sep 2, 2026 — _rdgao · 2026-08-29
- NTU releases ACE-Data-0: 150+ hours of time-synced multimodal embodied data — liuziwei7 · 2026-08-29
- KAIST Unveils K-Fold, Claiming 25x Faster Inference Than Google's AlphaFold — AllThingsApx · 2026-08-29
- Decathlon Switches to Chronos-2 for Large-Scale Demand Forecasting on AWS — AWS ML Blog · 2026-08-29
- Viewpoint: Optimizing for language compression led to current AI level — kuchaev · 2026-08-29