Tetris RL experiment: pretraining caps what RL can reach — PPO can provably converge to a bad policy
shizhediao · x · 2026-10-10
The author spent a week trying to move from shaped to terminal rewards on Tetris (backplay, prefix/continuation, credit assignment, randomness control) but stayed stuck under 150 lines, vs 550 with shaped rewards at the same model size. What worked: more pretraining.
- Key finding: pretraining largely determines where RL performance starts and levels off — it can permanently cap the model at a lower bound.
- With a Lean-certified proof (via ChatGPT), exact PPO can converge to a bad policy even when the model is capable of far better; a perfect critic and more iterations don't fix it.
- Two fixes: better pretraining, or process rewards.
- Implication: impressive "more RL iterations" results on LLMs may have a hidden caveat.
More from Research
- New kissing number lower bounds for dimensions 19-31, with C++ exact verification — felpix_ · 2026-10-10
- Mathematicians push back on 'solved problems = progress' as AI tests the metric — _onionesque · 2026-10-10
- BudgetPix: Google and UIUC's pixel diffusion model adapts compute per image, cutting tokens to 10% — CSProfKGD · 2026-10-10
- Anthropic Launches Conceptual Reasoning Fellows Program for Cog Sci PhDs — AndrewLampinen · 2026-10-10
- COLM Best Paper: Int-Bench Shows LLM Assistants Intervene Too Early and Too Often — maxhkw · 2026-10-10
- Demis Hassabis: Improving medicine is the most important thing AI can do, via new CZI partnership — DeryaTR_ · 2026-10-10