Dietterich: pretraining's next-token task already forces models to abstract and generalize
tdietterich · x · 2026-09-25
Pushing back on the "pretraining is just parroting" framing, AI researcher Thomas Dietterich argues the next-token task is demanding enough to force models to abstract and generalize in useful ways even before fine-tuning. SFT and RL then go beyond predicting language to predicting correct multi-token answers; pretraining mainly builds good internal representations plus a stochastic generator of candidate solutions.
Related event: Dietterich: LLMs Go Beyond Next-Token Prediction(2 posts)→
More from Research
- Protein design team secures tons of GPUs and lab equipment in single-day buildout — nlarusstone · 2026-09-25
- FD-Loss Achieves 0.75 One-Step Pixel-Space FID, Accepted as NeurIPS Oral — yuewang314 · 2026-09-25
- The 'No Magic OOD' Principle: Flagging a Common Flaw in Alignment Plans — nabla_theta · 2026-09-25
- NeurIPS 2026 paper shows noise-conditioning-free diffusion models implicitly implement Riemannian gradient flow — docmilanfar · 2026-09-25
- Solving the AI Science Validation Bottleneck: Why Dario's Gene-Editing Claim Isn't the Whole Story — Ghost_Pilot_MD · 2026-09-25
- AI PhD Laments: ICLR Submissions Are Formulaic While Industry Labs Sprint Ahead — vykthur · 2026-09-25