Tom Dietterich: SFT and RL go beyond next-token prediction in LLMs

tdietterich · x · 2026-09-25

Amid the debate over whether LLMs are "just predicting language," veteran ML researcher Thomas Dietterich explained that supervised fine-tuning and reinforcement learning go beyond single-token prediction to predict correct multi-token answers. Pre-training's next-token objective mainly serves to learn good internal representations and a stochastic generator of candidate solutions — his technical rebuttal to the "humans are just LLMs" framing.

Related event: Dietterich: LLMs Go Beyond Next-Token Prediction(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →