Researchers debate: does post-training stop a transformer from being a language model?

burny_tech · x · 2026-09-27

Responding to Melanie Mitchell's critique of LLMs, researcher francip argues that pretraining, SFT, preference optimization and RL all modify the same weights and the same inference process producing a conditional distribution over tokens. Post-training can radically change that distribution and behavior, but that doesn't make the result "not a language model" — just one optimized for a different objective. Part of the ongoing debate over whether modern LLMs are still next-token predictors.

Related event: Melanie Mitchell Sparks Fresh Debate Over "Stochastic Parrots" and What Counts as an LLM(28 posts)→

Original post →

More from Models

Models channel →