Tetris RL experiment: pretraining caps what RL can reach — PPO can provably converge to a bad policy

shizhediao · x · 2026-10-10

The author spent a week trying to move from shaped to terminal rewards on Tetris (backplay, prefix/continuation, credit assignment, randomness control) but stayed stuck under 150 lines, vs 550 with shaped rewards at the same model size. What worked: more pretraining.

Original post →

More from Research

Research channel →