Fireworks shows numerical mismatch can collapse RL training in 25 steps on GLM and MoE models
sophiamyang · x · 2026-10-02
Fireworks' Sophia Yang shared key findings on RL training infrastructure:
- Numerical mismatch can derail training: in a GLM 5.2 experiment, the same algorithm and data produced collapsing reward without alignment but stable reward over 25 steps.
- MoE complicates alignment: in Qwen3.5-MoE work, differences in how expert outputs were combined caused disagreement even when one implementation used higher precision.
- These discrepancies distort training updates and mimic data, reward, or learning-rate problems, leading teams into expensive dead-end debugging.
- Frontier training requires co-optimized training and inference: Fireworks develops and validates the trainer and rollout engine together across numerics, kernels, and MoEs.
The quoted Fireworks post notes rollouts drive most of RL's compute cost, and splitting rollout and training across engines is the root risk.
More from Infra
- Lambda closes $1B-plus GPU debt financing at 6.78% fixed rate, investment-grade rated — TheZachMueller · 2026-10-02
- Dual DGX Spark cluster vs M5 Ultra Mac Studio benchmark video is out — AIFlow_ML · 2026-10-02
- SageAttention up to 1.96x faster causal attention on RX 9070 XT via hand-written HIP fp8 kernel — Familiar_Worry332 · 2026-10-02
- Local GLM setup reportedly hits 175 tok/s with 98% draft acceptance on dual RTX 6000 Pros — HankYeomans · 2026-10-02
- Cloudflare lets non-admin members self-serve create Account API tokens via CLI, API and Terraform — irvinebroque · 2026-10-02
- SpaceX launches Google AI chips into orbit in push toward space-based data centers — pstAsiatech · 2026-10-02