Fireworks shows numerical mismatch can collapse RL training in 25 steps on GLM and MoE models

sophiamyang · x · 2026-10-02

Fireworks' Sophia Yang shared key findings on RL training infrastructure:

The quoted Fireworks post notes rollouts drive most of RL's compute cost, and splitting rollout and training across engines is the root risk.

Original post →

More from Infra

Infra channel →