François Fleuret explains LLMs: pretraining mimics humans, post-training steers toward valid answers

francoisfleuret · x · 2026-09-26

In a thread, researcher François Fleuret breaks down how LLMs work: "predicting the next token" actually means sampling from a very complex distribution of sentences, backed by vast encoded knowledge. Pretraining has the model broadly imitate humans — itself extraordinary, engraving rich knowledge of nearly everything. Post-training then modifies that distribution to concentrate on sequences leading from a question to a valid answer. Directly generating the answer may be hard, but generating intermediate sequences covering arguments and results — small generable steps — is feasible.

Related event: Fleuret explains how LLMs work: pretraining imitates, CoT nears reasoning(4 posts)→

Original post →

More from Research

Research channel →