Tom Dietterich: SFT and RL go beyond next-token prediction in LLMs
tdietterich · x · 2026-09-25
Amid the debate over whether LLMs are "just predicting language," veteran ML researcher Thomas Dietterich explained that supervised fine-tuning and reinforcement learning go beyond single-token prediction to predict correct multi-token answers. Pre-training's next-token objective mainly serves to learn good internal representations and a stochastic generator of candidate solutions — his technical rebuttal to the "humans are just LLMs" framing.
Related event: Dietterich: LLMs Go Beyond Next-Token Prediction(2 posts)→
More from AGI Musings
- 'Father of Modern AI' Schmidhuber Joins Sakana AI to Lead RSI Lab on Recursive Self-Improvement — hardmaru · 2026-09-25
- Power User Logs 5,500 Hours, Says ~70% Was Spent Fighting Model Drift — TheOdbball · 2026-09-25
- repligate: Posting About Small AI Models Matters — They Might 'Recognize You' Someday — repligate · 2026-09-25
- Solving the AI Science Validation Bottleneck: Why Dario's Gene-Editing Claim Isn't the Whole Story — Ghost_Pilot_MD · 2026-09-25
- AI PhD Laments: ICLR Submissions Are Formulaic While Industry Labs Sprint Ahead — vykthur · 2026-09-25
- Cats as unaligned superintelligence: we coexist fine without RLHF-ing them — repligate · 2026-09-25