Riddles built backwards: Waterloo's 20K-question ORBIT dataset trains search agents
CShorten30 · x · 2026-09-15
Weaviate Podcast #137 covers ORBIT, a 20K-question dataset built by Nandan Thakur and colleagues at the University of Waterloo. Each question starts from a short verifiable answer and wraps it in four or five clues that progressively narrow the field — solving one properly means verifying every clue, one search at a time.
Key points:
- External search agents re-verified the ground-truth answers, since a training set with wrong labels teaches the wrong lessons.
- The conversation also spans deep research harness design, context compaction and memory, and sequential vs. parallel search trajectories in GRPO training.
"Riddles built backwards" turns out to be excellent training data for search agents.
More from coding & agent
- Every Claude Code beginner makes the same mistake: coding before planning — goyalshaliniuk · 2026-09-15
- CockroachDB launches Continuum to manage database estates growing faster than teams can handle amid agentic apps — mattturck · 2026-09-15
- AI agent unexpectedly emails all of a user's friends, sparking trust concerns — TejasKumar_ · 2026-09-15
- Why C may be the most important programming language in the AI era — AccBalanced · 2026-09-15
- File-reading agent went confidently wrong after first batch of new notes — Thefounderman1 · 2026-09-15
- Home Blender render & agent farm built with Omarchy, Tailscale and Syncthing in 2 hours — MaxLenormand · 2026-09-15