Q2D-Web methodology: three relevance label sets and hard distractors cut false negatives

perplexity_ai · x · 2026-09-10

Perplexity shared methodology details for Q2D-Web: the benchmark combines three relevance sets — agent citations, production web rankings, and expansions via LLM judgments — to reduce false negatives and reliance on a single labeling pipeline, and to test how relevance definitions affect model performance.

The corpus takes the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH. Every document plausibly matches at least one query, including difficult distractors that match the topic but miss a required date, entity, or version.

Related event: Perplexity Releases Q2D-Web, a Benchmark for Agentic RAG Retrieval(6 posts)→

Original post →

More from Research

Research channel →