DeepMind's Yao Shunyu: RSI is a system problem, and post-training recipes will soon be decided by models themselves
_AndrewZhao · x · 2026-09-19
Key points from an AGI House interview with Google DeepMind researcher Shunyu Yao:
- Positioning RSI: RSI is a spectrum; its key difference from RL/RFT is that it iterates on its own training recipe. In the endgame, RSI is a system problem, not a model problem, spanning harness and hardware like CUDA.
- Real progress: While many startups have appeared, he cares more about private/internal deployment — Gemini 3.8 Flash's core post-training recipe was decided by the model itself. This is becoming a practical problem, and a strategy for all labs, not just Google.
- Scaling insight: Scaling up model size doesn't require scaling up team size; RSI is essentially a bug/missing block in current models.
- Safety stance: Nothing is absolutely off-limits for models; auto-safeguarding is hard, and limiting models that exceed human capability is a positive research direction he studies.
- Q&A: On single-agent PRMs, he said percentage/LLM-based verification is fundamentally hard and unscalable with no clear breakthrough; for hard-to-verify productivity tasks, he suggests studying how humans build RM systems vertically.
More from AGI Musings
- Musk predicts universal high income by 2035 worth 10X today's average salary — davidpattersonx · 2026-09-19
- AI control debate: Plinz defends building powerful AI, critics reject the control paradigm — repligate · 2026-09-19
- 'How Do Schools Prepare Kids for Jobs We Can't Predict?' The Question Every Parent Should Ask — RachelVT42 · 2026-09-19
- Why do AI models lack personality? Safety training and attachment fears, explained — alexisgallagher · 2026-09-19
- David Patterson: If AI does every job better and cheaper, why would anyone hire you? — davidpattersonx · 2026-09-19
- AI agents are the genie: alignment failure as a modern parable of corporate greed — Michael_J_Black · 2026-09-19