Q2D-Web benchmark debuts: 70K agent queries to evaluate retrievers across 190M web docs
antoine_chaffin · x · 2026-09-12
A new benchmark and public leaderboard, Q2D-Web, targets first-stage retrievers for agentic search at production scale.
- 70,000 agent queries, 190 million web documents, and three sets of relevance judgements
- Designed to measure how well retrievers feed agent search pipelines
- Public leaderboard, eval request form, and research blog are live
More from Research
- Columbia to host symposium on AI, genomics and community science in biogeography — sarameghanbeery · 2026-09-12
- Japan's JST-MEXT to host AI for Science 2026 symposium featuring Google DeepMind — heiga_zen · 2026-09-12
- Fruit Fly Brain Simulation Solves Rubik's Cube, Igniting Consciousness Debate — sebkrier · 2026-09-12
- 25 Fields Medal winners warn AI's goals are 'severely misaligned' with mathematics — The Decoder · 2026-09-12
- LuxoBench: A New AI Benchmark Tasks Models With Building Real Electromechanical Devices — vincent_koc · 2026-09-12
- Coding Is Not All You Need: CMU author argues GPT-6's robot tasks hit a world-model wall — ceciletamura · 2026-09-12