BrowseComp-Plus benchmarks deep-research agents and highlights 9-turn production systems

IgorCarron · x · 2026-07-22

BrowseComp-Plus introduces a fair benchmark for deep-research agents

A new benchmark called BrowseComp-Plus is presented as a fair, disentangled evaluation for deep-research systems.

It is built on top of BrowseComp and uses a carefully curated corpus of web documents with human-verified positives and mined hard negatives.

The benchmark is designed to separately evaluate:

The screenshot shows LightOn Console reaching the top 5 in the production vs. research comparison, with about 9 turns on average versus 209 turns for the top-ranked research entry in one setting, highlighting big efficiency differences.

The broader message is that production systems can be both competitive and far cheaper to run than leaderboard-topping research experiments.

Original post →

More from Research

Research channel →