NVIDIA Data Designer workshop: 50k seeds become 7k records teaching LLMs to search

AI Engineer · youtube · 2026-10-12

An AI Engineer workshop by NVIDIA's Dhruv Nathawani shows how to build a synthetic data pipeline with Data Designer that teaches models to search instead of answering from memory.

Pipeline

Three staged notebooks: (1) Q&A generation with seed topics, difficulty samplers, prompt templates, structured outputs and model judges — preview records before batching; (2) tool access and search trace capture; (3) end-to-end riddle generation with full agent trajectories.

Key lessons: verifying answers is hard when source graphs go stale, so combine model judgments, deterministic checks and human review; use synthetic personas to preserve demographic distributions. Design for diversity, inspect examples, overproduce, filter, and version the recipe.

Original post →

More from coding & agent

coding & agent channel →