Mistral details RL at scale: 3k GPUs, ~33B tokens/day, stable long-horizon training

sophiamyang · x · 2026-10-06

Mistral researcher Sophia Yang shared key design points of the company's large-scale reinforcement learning training stack:

In a reply she also claimed Mistral is "crushing GLM 5.3 on human evals across domains" — unverified, with no detailed data attached.

Original post →

More from Models

Models channel →