Perplexity Releases Q2D-Web: 190M-Doc Benchmark for Agentic RAG Retrieval
_reachsumit · x · 2026-09-09
Perplexity AI introduced Q2D-Web, a large-scale benchmark for first-stage retrieval in agentic RAG, pairing a 190M-document web corpus with 70k agent-reformulated queries drawn from production user queries across ten languages. It addresses gaps in existing benchmarks, which either have huge corpora but few queries, or many queries but small corpora, and typically test human-written queries rather than machine reformulations. Three relevance judgment sets are provided (agent citations, production rankings, and a union augmented with LLM judgments). Benchmarks across 13 lexical, dense, and late-interaction retrievers show rankings are largely insensitive to judgment-set choice.
More from Research
- Timothy Duff's ECCV 2026 SfM-DL workshop slides on algebraic optimality for minimal solvers — ducha_aiki · 2026-09-09
- Drop a fixed batch proportion instead of per-sample tokens: capi author shares training trick — giffmana · 2026-09-09
- Adding Greek to a Cosmos3 VLA policy: bilingual training helps but lags far behind English — KIEFERSA · 2026-09-09
- Transformers encode a partner's expertise early but only act on it in later layers — Mika Okamoto · 2026-09-09
- Cadence uses a time-series foundation model for error-bounded lossy compression of demand data — Roberto Tacconelli · 2026-09-09
- Fourth Perception Test Challenge at ECCV 2026 pushes multimodal models on city-scale spatial intelligence — AjdDavison · 2026-09-09