Perplexity launches Q2D-Web, a benchmark and leaderboard for agentic RAG retrieval
perplexity_ai · x · 2026-09-10
Perplexity released Q2D-Web (Query2Doc-Web), a public benchmark and leaderboard evaluating retrieval in agentic RAG systems, testing how embedding models handle agent-reformulated search queries at web scale.
Key details:
- Corpus: top 5,000 production retrieval results per query, deduplicated with MinHash-LSH; every document plausibly matches at least one query, including hard distractors that match the topic but miss a required date, entity, or version.
- Relevance labels: three sets built from agent citations, production web rankings, and LLM-judgment expansions, reducing false negatives and single-pipeline dependence.
- Coverage: programming, law, health, science, finance, consumer goods, travel, entertainment, and local info across ten languages (English 65.8%).
Related event: Perplexity Releases Q2D-Web, a Benchmark for Agentic RAG Retrieval(6 posts)→
More from Research
- FrogNano: a 4B model trained purely with RL on synthetic tasks hits repo-level coding — burkov · 2026-09-10
- Stanford lab rebuilt as interactive 3D web scene in a day with GPT-6 Astra and Retriever — OfirPress · 2026-09-10
- VisionCoach: RL framework rewards correct visual attention for grounded video reasoning, SOTA zero-shot — mohitban47 · 2026-09-10
- Goodfire explains how probes can read model minds to catch cyber intent and reward hacking — leland_mcinnes · 2026-09-10
- Engineer deploys hundreds of parallel AI agents to work on a type 1 diabetes cure — Scobleizer · 2026-09-10
- SOFAIR lab launches with UCL, Cambridge, Oxford, Edinburgh to do Science for AI — latticecut · 2026-09-10