BrowseComp-Plus: New Benchmark for Deep-Research Agent Evaluation

lintool · x · 2026-08-22

Castorini introduced BrowseComp-Plus, a benchmark for evaluating deep-research agents. Built on the 553M-document ClimbMix-400B corpus, it features human-verified positives and web-mined hard negatives. It aims to provide a fair evaluation for components in deep-research systems. Code and datasets are available on GitHub.

Original post →

More from Research

Research channel →