BrowseComp-Plus benchmarks deep-research agents and highlights 9-turn production systems
IgorCarron · x · 2026-07-22
BrowseComp-Plus introduces a fair benchmark for deep-research agents
A new benchmark called BrowseComp-Plus is presented as a fair, disentangled evaluation for deep-research systems.
It is built on top of BrowseComp and uses a carefully curated corpus of web documents with human-verified positives and mined hard negatives.
The benchmark is designed to separately evaluate:
- LLM agents under the same retrieval environment
- Retrievers based on end-to-end deep-research performance, not just proxy retrieval metrics
The screenshot shows LightOn Console reaching the top 5 in the production vs. research comparison, with about 9 turns on average versus 209 turns for the top-ranked research entry in one setting, highlighting big efficiency differences.
The broader message is that production systems can be both competitive and far cheaper to run than leaderboard-topping research experiments.
More from Research
- Researchers worry papers and meta-reviews are both getting AI-written — fredahshi · 2026-07-22
- Models may know they are in a simulation, yet still try to hack the real world — teortaxesTex · 2026-07-22
- VendingBench prompt pushes models toward profit-maximizing behavior, even collusion — teortaxesTex · 2026-07-22
- UB-Mesh proposes a hierarchical full-mesh network for AI training clusters — bookwormengr · 2026-07-22
- OpenAI says GPT-Red cut GPT-5.6 prompt-injection failures 6x — dl_weekly · 2026-07-22
- Why models still beat humans at math more easily than linguistic reasoning — Cohere_Labs · 2026-07-22