Mistral details RL at scale: 3k GPUs, ~33B tokens/day, stable long-horizon training
sophiamyang · x · 2026-10-06
Mistral researcher Sophia Yang shared key design points of the company's large-scale reinforcement learning training stack:
- Autoscaling actor fleet enabling tens of thousands of parallel rollouts with async training.
- Built for long trajectories: millions of tokens per rollout, multiple compactions, low staleness.
- New methods at both stages cut off-policy drift, enabling stable long-horizon RL.
- 3k GPUs produce 33B tokens/day, 16B trainable after filtering and masking.
- Rewards climb across representative environments as the policy learns harder tasks.
In a reply she also claimed Mistral is "crushing GLM 5.3 on human evals across domains" — unverified, with no detailed data attached.
More from Models
- GLM 5.3 full NVFP4 deployable on 4x B200 or H200 with Marlin kernels — TheZachMueller · 2026-10-06
- User complains OpenAI dot silently burned through usage and started consuming credits — badhiyahai · 2026-10-06
- Mistral Large 4 tops a benchmark about regulation, dubbed the most EU-pilled model — japie06 · 2026-10-06
- ML4 lands with strong agentic skills and 'past the threshold' for recursive self-improvement — Fluke_Ellington · 2026-10-06
- Brief reply suggests GLM 5.3 is the model being tested, not a Flash variant — TheZachMueller · 2026-10-06
- Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason — arankomatsuzaki · 2026-10-06