DeepMind's Nathaniel Daw: the post-training critique has held since ChatGPT

burny_tech · x · 2026-09-27

In the ongoing debate with Melanie Mitchell, DeepMind researcher Nathaniel Daw notes that the shift away from pure next-token prediction via instruction tuning and RLHF has been true for years, at least since ChatGPT — slightly after the stochastic parrots paper, whose authors never retracted their core critique in light of those techniques.

Related event: Melanie Mitchell Sparks Fresh Debate Over "Stochastic Parrots" and What Counts as an LLM(28 posts)→

Original post →

More from Models

Models channel →