François Fleuret explains LLMs: pretraining mimics humans, post-training steers toward valid answers
francoisfleuret · x · 2026-09-26
In a thread, researcher François Fleuret breaks down how LLMs work: "predicting the next token" actually means sampling from a very complex distribution of sentences, backed by vast encoded knowledge. Pretraining has the model broadly imitate humans — itself extraordinary, engraving rich knowledge of nearly everything. Post-training then modifies that distribution to concentrate on sequences leading from a question to a valid answer. Directly generating the answer may be hard, but generating intermediate sequences covering arguments and results — small generable steps — is feasible.
Related event: Fleuret explains how LLMs work: pretraining imitates, CoT nears reasoning(4 posts)→
More from Research
- Anthropic says Claude can compute notoriously hard Nine Loops physics amplitudes — daniel_mac8 · 2026-09-26
- Researchers call 'next token predictor' a contentless critique of LLMs — aran_nayebi · 2026-09-26
- Researchers teach LLMs to find interesting theorems, boosting interestingness 4.3x — CatAstro_Piyush · 2026-09-26
- Anthropic: Claude solves nine-loop scattering amplitudes, breaking the eight-loop record — AnthropicAI · 2026-09-26
- Both "grep is all you need" and "BM25 is all you need" Papers Just Got Accepted — lintool · 2026-09-26
- Reasonable Team Publishes TLA+ Tutorial: Not a Silver Bullet, AI Agents Could Change That — fhuszar · 2026-09-26