François Fleuret's Thread: LLMs Are More Than 'Next-Token Predictors'
On September 26, Meta researcher François Fleuret posted a seven-tweet thread on Twitter systematically responding to the popular dismissive claim that "LLMs just predict the next token." His core conclusion: the description is literally true but operates at entirely the wrong level of abstraction and carries zero information.
Confirmed
- Fleuret argued that saying a model "just predicts the next token" is like saying a chess program "just moves electrons": true as a description, yet not the right level at which to understand the system—LLMs should be evaluated at the abstraction level of sentence distributions.
- What's called "predicting the next token" is actually sampling from an extremely complex sentence distribution; just as a person can predict where billiard balls will go without looking at electrons, the model supports this distribution with vast knowledge.
- In pre-training, the model "just" imitates humans broadly, but this step is itself crucial because it engraves rich knowledge about nearly everything into the model.
- In post-training, the model learns to modify this distribution to focus on generating valid answers.
- On reasoning: a model may struggle to directly produce an answer in one step, but can perfectly well generate sequences of tokens covering intermediate arguments and intermediate results—a series of small steps generated one at a time. Fleuret argued that the real meaning of "it just predicts the next word" thus becomes "given a question, generate incremental small steps," and this chain-of-thought-style generation of intermediate results comes close to genuine reasoning.
Why it matters
- "Just predicting the next token" is the most commonly cited argument for dismissing LLM capabilities, and Fleuret's breakdown provides a clear hierarchical framework for rebuttal: mechanistic description and capability assessment should be kept at their appropriate levels of abstraction.
- His clear division of labor between pre-training (knowledge injection) and post-training (distribution correction) helps the public understand why the seemingly simple goal of "imitating humans" yields broad knowledge, and how post-training steers models toward valid answers.
- His reading of chain-of-thought—approximating reasoning through incrementally generated small steps—takes a leaning-toward-yes stance in the debate over whether LLMs truly reason.
2026-09-26 ~ 2026-09-26 · 6 related posts
Primary sources
- François Fleuret: 'Just Predicting the Next Token' Is the Wrong Level of Description — francoisfleuret ·
- François Fleuret explains LLMs: pretraining mimics humans, post-training steers toward valid answers — francoisfleuret ·
- Fleuret: Generating Intermediate Steps Is 'Very Close to Actual Reasoning' — francoisfleuret ·
- 'Just Moving Electrons': Fleuret's Chess Analogy Against LLM Reductionism — francoisfleuret · 2026-09-26
- [source] François Fleuret: 'Just Predicting the Next Token' Is the Wrong Level of Description — francoisfleuret · 2026-09-26
- [source] François Fleuret explains LLMs: pretraining mimics humans, post-training steers toward valid answers — francoisfleuret · 2026-09-26
- [source] Fleuret: Generating Intermediate Steps Is 'Very Close to Actual Reasoning' — francoisfleuret · 2026-09-26
2 near-duplicate retellings: francoisfleuret · francoisfleuret