Q2D-Web methodology: three relevance label sets and hard distractors cut false negatives
perplexity_ai · x · 2026-09-10
Perplexity shared methodology details for Q2D-Web: the benchmark combines three relevance sets — agent citations, production web rankings, and expansions via LLM judgments — to reduce false negatives and reliance on a single labeling pipeline, and to test how relevance definitions affect model performance.
The corpus takes the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH. Every document plausibly matches at least one query, including difficult distractors that match the topic but miss a required date, entity, or version.
Related event: Perplexity Releases Q2D-Web, a Benchmark for Agentic RAG Retrieval(6 posts)→
More from Research
- CoopEval: a framework for comparing cooperation mechanisms in multi-agent systems — conitzer · 2026-09-10
- Open Yap 1K: 1,000 hours of natural two-speaker conversations, free for commercial use — realmrfakename · 2026-09-10
- Hank Yang: AI Excels at Well-Defined Problems, So the Real Skill Is Defining New Ones — hankyang94 · 2026-09-10
- Jacobian conjecture drama: Anthropic's Alpoge responds to leaked BGV paper concerns — suchenzang · 2026-09-10
- NNsight 0.8 pre-release ships faster engine, MoE and near-native vLLM support — davidbau · 2026-09-10
- ECCV talk outlines three pillars for embodied AI: motion prediction, evidence, streaming — CSProfKGD · 2026-09-10