BrowseComp-Plus: New Benchmark for Deep-Research Agent Evaluation
lintool · x · 2026-08-22
Castorini introduced BrowseComp-Plus, a benchmark for evaluating deep-research agents. Built on the 553M-document ClimbMix-400B corpus, it features human-verified positives and web-mined hard negatives. It aims to provide a fair evaluation for components in deep-research systems. Code and datasets are available on GitHub.
More from Research
- Equilibrium Forcing: Adaptive video generation without noise conditioning — kwangmoo_yi · 2026-08-22
- Equilibrium Forcing Enables Adaptive Video Generation Without Noise Conditioning — kwangmoo_yi · 2026-08-22
- Deep Dive: How World Models Unlock Physical AI and Robotics — risingodegua · 2026-08-22
- OpenBind Released: Small Molecule Co-folding Model Based on OpenFold3 — MoAlQuraishi · 2026-08-22
- Research Uses LLM Code Gen to Drive Evolutionary Algorithms — kenneth0stanley · 2026-08-22
- Yukon Leaderboard Updated to Rank by Total Speedup Contribution — morgymcg · 2026-08-22