Perplexity launches Q2D-Web, a 190M-doc benchmark and leaderboard for agentic RAG retrieval
perplexity_ai · x · 2026-09-10
Perplexity introduced Q2D-Web (Query2D-Web), a large-scale benchmark and public leaderboard for evaluating embedding models on retrieval in agentic RAG systems.
- Built from 23,000 PII-free production searches collected over nine months, spanning ten languages (65.8% English) and dozens of domains from programming and law to health, science, finance, and travel.
- Agents reformulate user requests into primary and support queries, each evaluated independently with its own relevance judgments.
- The combined set includes 190M web documents, 69,721 agent-reformulated queries, and an average of 99.6 positive judgments per query to reduce false negatives.
- Three relevance sets (agent citations, production web rankings, LLM-expanded judgments) reduce reliance on a single labeling pipeline.
- 13 retrieval models evaluated with Recall@1000: pplx-embed-v1-4b leads Web Ranking (65.73) and Combined (69.11); Nemotron-3-Embed-8B leads Citation (61.68).
- RRF-based subsampling preserves full-corpus rankings using only 31.7% of documents, cutting pplx-embed-v1-4b evaluation from 4,608 to 1,500 H200 GPU-hours.
Related event: Perplexity Releases Q2D-Web, a Benchmark for Agentic RAG Retrieval(6 posts)→
More from Infra
- NVIDIA joins the Rust Foundation — blelbach · 2026-09-10
- Epoch estimates OpenAI quadrupled compute in both 2024 and 2025, a 17x two-year jump — FlorianGallwitz · 2026-09-10
- Hands-On Guide: Safely Running Untrusted Code with Google Cloud Run Sandboxes — rseroter · 2026-09-10
- Nvidia NVL72 rack shipments seen up 50% in 2027, output forecast to top $710B — Beth_Kindig · 2026-09-10
- Stealth 7-year startup Kepler debuts AI memory beyond HBM, secures up to $245M in US government support — npinto · 2026-09-10
- Radium racked its own GPUs to undercut OpenAI and Anthropic pricing — small devs still won't switch — No_Raspberry7273 · 2026-09-10