NVIDIA Data Designer workshop: 50k seeds become 7k records teaching LLMs to search
AI Engineer · youtube · 2026-10-12
An AI Engineer workshop by NVIDIA's Dhruv Nathawani shows how to build a synthetic data pipeline with Data Designer that teaches models to search instead of answering from memory.
Pipeline
- Wikidata entity paths become multi-hop "search riddles"; 50,000 seed examples funnel down to 7,000 filtered training records.
- Tavily provides search via MCP; the output includes full question, tool-call and response trajectories for SFT/distillation, with RL alternatives discussed.
Three staged notebooks: (1) Q&A generation with seed topics, difficulty samplers, prompt templates, structured outputs and model judges — preview records before batching; (2) tool access and search trace capture; (3) end-to-end riddle generation with full agent trajectories.
Key lessons: verifying answers is hard when source graphs go stale, so combine model judgments, deterministic checks and human review; use synthetic personas to preserve demographic distributions. Design for diversity, inspect examples, overproduce, filter, and version the recipe.
More from coding & agent
- AI-era coding interview: hand candidates a broken agent-built app and watch them debug — gethackteam · 2026-10-12
- Visualizing parallel requests with TQDM: watch multiple queues grow in real time — AnuranBuilds · 2026-10-12
- Running all your coding agents through a Grok bot: parallel Claude Code PR pipeline — Rasmic · 2026-10-12
- Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context — matteiuspi · 2026-10-12
- plain writing is the team's single most-used internal skill, by a factor of 2 — sh_reya · 2026-10-12
- plain-writing-skill: an open-source skill that makes AI agents write plainly — sh_reya · 2026-10-12