DeepMind researcher: add RL rollouts to pretraining and nothing fundamentally changes

burny_tech · x · 2026-09-27

DeepMind researcher Tomasz Korbak debates Melanie Mitchell on whether RL post-training moves LLMs beyond language modeling. His take: it's still modeling language from a slightly different distribution — the architecture and objective (for a fixed dataset) are basically the same, and adding RL rollouts to pretraining data wouldn't change the effect. He also argues the stochastic parrots paper made claims about the whole next-token paradigm without carving out this distinction.

Original post →

More from Models

Models channel →