Perplexity releases Q2D-Web: a 190M-document retrieval benchmark for agentic RAG
antoine_chaffin · x · 2026-09-10
Perplexity and collaborators released Q2D-Web, a large-scale retrieval benchmark and public leaderboard for agentic RAG systems (arXiv preprint).
- 190M-document web corpus paired with 70K agent-reformulated queries in 10 languages, derived from real production user queries.
- Motivation: existing benchmarks either have huge corpora but few queries, or many queries but only million-scale corpora — and none test machine-reformulated queries, whose distribution differs from human search.
- Three relevance judgment sets: agent citations, production rankings, and a union plus LLM-based judgments to reduce false negatives.
- Benchmarked 13 retrievers (lexical, dense, late-interaction); relative ordering is largely insensitive to judgment-set choice.
- Public leaderboard included.
More from Research
- Meta's Boxer at ECCV: closing the 3D ground-truth gap with 2D scaling — ducha_aiki · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- SyncWorld turns world models into zero-shot robot simulators via visual calibration — Yuncong Yang · 2026-09-10
- Feng Yao wins ECVA PhD Award at ECCV 2026 for 3D humans + language thesis — Michael_J_Black · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10
- MIT's Point2Pose tracks unknown objects in 6D with full-occlusion recovery — joemeno · 2026-09-10