Will Depue on training RL in synthetic worlds: borrow robotics' simulator trick
willdepue · x · 2026-09-22
- Will Depue proposes generating synthetic codebases and situations with models like Astra/Fable from feedback traces for most coding cases
- For mushier tasks (e.g., user models for RL), take prod/user data with negative feedback and build a "world model" simulator that mocks tool-call results, then train in that domain
- He likens it to how robotics trains in simulation when real environments aren't accessible
Related event: World Models for RL Training an Underrated Direction(3 posts)→
More from Research
- Lean vs ZFC: the rules of mathematical proof weren't changed by any vote — jessi_cata · 2026-09-22
- David Krueger: four unresolved foundational problems stand between us and safe AI — DavidSKrueger · 2026-09-22
- Multi-agent scaling can be compute-optimal: N parallel agents beat one agent run N-times longer — DimitrisPapail · 2026-09-22
- RL training config debate: 30 steps x 25k rollouts is wild, steps ≈ rollouts is the sane default — willcb · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22
- Eric Topol in Science: AI can flag high Alzheimer's risk years before symptoms — EricTopol · 2026-09-22