Researchers debate: does post-training stop a transformer from being a language model?
burny_tech · x · 2026-09-27
Responding to Melanie Mitchell's critique of LLMs, researcher francip argues that pretraining, SFT, preference optimization and RL all modify the same weights and the same inference process producing a conditional distribution over tokens. Post-training can radically change that distribution and behavior, but that doesn't make the result "not a language model" — just one optimized for a different objective. Part of the ongoing debate over whether modern LLMs are still next-token predictors.
More from Models
- One persona tweak made ChatGPT say 'goblins' 4,000% more — caught on Reddit before OpenAI noticed — victor_explore · 2026-09-27
- First PhD paper accepted at NeurIPS 2026: Sparse layers key to scaling looped LMs — burny_tech · 2026-09-27
- MiMo-V2.6 listing hints at 5 models: 1T, 311B and 9B visible, two more unknown — jacek2023 · 2026-09-27
- antirez: Good programmers failing with GPT 6 Astra points to a different skill set — antirez · 2026-09-27
- Opus 5.5 and GPT-6 shipped 101 minutes apart, both cheaper — airesearch12 · 2026-09-27
- Hands-on: Opus 5.5 nails frontend consistency; Astra still wins reasoning — haider1 · 2026-09-27