Mistral researcher: 15T tokens underestimate pretraining, closer to 40T+

eliebakouch · x · 2026-10-08

Former Mistral researcher Elie Bakouch argues the commonly cited 15T token estimate for pretraining is too low — the real figure is likely 40T+ tokens, which would multiply wall-clock training time by 3-4x. He finds the assumed MFU/goodput reasonable or slightly ambitious.

He also stresses a common confusion: pretraining FLOPs should be computed from active parameters, not total parameters (citing the "a 1000-parameter model" claim as an example). The point isn't precision but showing the compute scale is achievable.

Original post →

More from Infra

Infra channel →