Speculation: Grok Pro line is an extension of Flash line, mxfp4 QAT likely speeds RL rollouts
stochasticchasm · x · 2026-09-22
A poster analyzes xAI's model lineup, arguing the Pro tier reads as an extension of the Flash line — a good sign for compute allocation. On training details, they note "no loss spike" doesn't necessarily rule out degraded performance, just lower eval scores, and speculate that mxfp4 QAT during midtraining is now standard practice and appears specifically aimed at accelerating RL rollouts. Unconfirmed speculation.
More from Models
- Award-Winning French Novel Detected as Nearly 100% AI-Written, Sparking Copyright Debate — TuhinChakr · 2026-09-22
- Xiaomi Releases Small Qwen 3.5 9B Distill SFT'd on MiMo Data, Plus RL Environments — teortaxesTex · 2026-09-22
- OpenAI has largely automated training experimental models; new model cracks 100 open math problems — kimmonismus · 2026-09-22
- Grok 4.7 scores 100% on music error detection test, matching GPT-6 Astra — yunta_tsai · 2026-09-22
- MiMo-V2.6 called one of the largest open-source RL runs, now top open model — andrew_n_carr · 2026-09-22
- Dev on AI coding's biggest pain point: models are still too stupid, slow and expensive — remilouf · 2026-09-22