Dietterich: pretraining's next-token task already forces models to abstract and generalize

tdietterich · x · 2026-09-25

Pushing back on the "pretraining is just parroting" framing, AI researcher Thomas Dietterich argues the next-token task is demanding enough to force models to abstract and generalize in useful ways even before fine-tuning. SFT and RL then go beyond predicting language to predicting correct multi-token answers; pretraining mainly builds good internal representations plus a stochastic generator of candidate solutions.

Related event: Dietterich: LLMs Go Beyond Next-Token Prediction(2 posts)→

Original post →

More from Research

Research channel →