Q2D-Web corpus: nine domains, ten languages, English at 65.8%
perplexity_ai · x · 2026-09-10
Perplexity detailed Q2D-Web's data composition: the corpus combines the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH. Each document plausibly matches at least one query, including hard distractors that match the topic but miss a required date, entity, or version.
Queries span programming, law, health, science, finance, consumer goods, travel, entertainment, and local information across ten languages, with English accounting for 65.8%.
More from Research
- ECCV talk outlines three pillars for embodied AI: motion prediction, evidence, streaming — CSProfKGD · 2026-09-10
- FrogNano: a 4B model trained purely with RL on synthetic tasks hits repo-level coding — burkov · 2026-09-10
- Stanford lab rebuilt as interactive 3D web scene in a day with GPT-6 Astra and Retriever — OfirPress · 2026-09-10
- VisionCoach: RL framework rewards correct visual attention for grounded video reasoning, SOTA zero-shot — mohitban47 · 2026-09-10
- Goodfire explains how probes can read model minds to catch cyber intent and reward hacking — leland_mcinnes · 2026-09-10
- Engineer deploys hundreds of parallel AI agents to work on a type 1 diabetes cure — Scobleizer · 2026-09-10