DeepMind's Nathaniel Daw: the post-training critique has held since ChatGPT
burny_tech · x · 2026-09-27
In the ongoing debate with Melanie Mitchell, DeepMind researcher Nathaniel Daw notes that the shift away from pure next-token prediction via instruction tuning and RLHF has been true for years, at least since ChatGPT — slightly after the stochastic parrots paper, whose authors never retracted their core critique in light of those techniques.
More from Models
- One persona tweak made ChatGPT say 'goblins' 4,000% more — caught on Reddit before OpenAI noticed — victor_explore · 2026-09-27
- First PhD paper accepted at NeurIPS 2026: Sparse layers key to scaling looped LMs — burny_tech · 2026-09-27
- MiMo-V2.6 listing hints at 5 models: 1T, 311B and 9B visible, two more unknown — jacek2023 · 2026-09-27
- antirez: Good programmers failing with GPT 6 Astra points to a different skill set — antirez · 2026-09-27
- Opus 5.5 and GPT-6 shipped 101 minutes apart, both cheaper — airesearch12 · 2026-09-27
- Hands-on: Opus 5.5 nails frontend consistency; Astra still wins reasoning — haider1 · 2026-09-27