DeepMind researcher: add RL rollouts to pretraining and nothing fundamentally changes
burny_tech · x · 2026-09-27
DeepMind researcher Tomasz Korbak debates Melanie Mitchell on whether RL post-training moves LLMs beyond language modeling. His take: it's still modeling language from a slightly different distribution — the architecture and objective (for a fixed dataset) are basically the same, and adding RL rollouts to pretraining data wouldn't change the effect. He also argues the stochastic parrots paper made claims about the whole next-token paradigm without carving out this distinction.
More from Models
- Limite 1B Violetto: compact Apache 2.0 model focused on math and reasoning — tensorqt · 2026-09-27
- User shows Opus 5.5 finishing a complex task in 15 minutes — gaganghotra_ · 2026-09-27
- Classic ROME Paper Revisited: GPT Stores Facts as MLP Key-Value Pairs You Can Edit — burny_tech · 2026-09-27
- JevBench researcher talks Jev-class models on ThursdAI podcast (from 2:01:00) — airesearch12 · 2026-09-27
- Jev pricing reality check: 720 states x 30 questions likely costs under 10 cents — schwentker · 2026-09-27
- Is RL-trained reasoning still next-token prediction? The loss function debate — burny_tech · 2026-09-27