Q2D-Web corpus: nine domains, ten languages, English at 65.8%

perplexity_ai · x · 2026-09-10

Perplexity detailed Q2D-Web's data composition: the corpus combines the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH. Each document plausibly matches at least one query, including hard distractors that match the topic but miss a required date, entity, or version.

Queries span programming, law, health, science, finance, consumer goods, travel, entertainment, and local information across ten languages, with English accounting for 65.8%.

Original post →

More from Research

Research channel →