Mistral researcher: 15T tokens underestimate pretraining, closer to 40T+
eliebakouch · x · 2026-10-08
Former Mistral researcher Elie Bakouch argues the commonly cited 15T token estimate for pretraining is too low — the real figure is likely 40T+ tokens, which would multiply wall-clock training time by 3-4x. He finds the assumed MFU/goodput reasonable or slightly ambitious.
He also stresses a common confusion: pretraining FLOPs should be computed from active parameters, not total parameters (citing the "a 1000-parameter model" claim as an example). The point isn't precision but showing the compute scale is achievable.
More from Infra
- CoreWeave ships a wave of products at Fully Connected event — thursdai_pod · 2026-10-08
- DuckDB shows a database with no data: a few hundred KB catalog over your data lake — RexDouglass · 2026-10-08
- Yudkowsky: pretraining is down to 7% of compute and spiky capability predictions are back — repligate · 2026-10-08
- National Compute goes live with hundreds of MI355x and B300 nodes, subsidized for .edu and .gov — ycombinator · 2026-10-08
- Microsoft demos local models at Windows launch, shows off GitHub Auto routing — DanWahlin · 2026-10-08
- xAI Bets on Leasing Data Centers Over Owning a Frontier Lab, Says Investor — PaulYacoubian · 2026-10-08