Fireworks: numerical mismatches can collapse RL reward training, MoE makes it worse

sophiamyang · x · 2026-10-02

Fireworks' Sophia Yang summarizes why training and rollout engines must be co-designed: in a GLM 5.2 experiment, identical algorithm and data produced collapsing reward without numerics alignment but stayed stable with it over 25 steps. A Qwen3.5-MoE investigation found expert-output combination differences routed tokens to different experts even at higher precision. These mismatches mimic data/reward/learning-rate bugs, sending teams through costly misdirected debugging. Fireworks develops and validates trainer and rollout engine together to keep RL scaling consistent.

Related event: Fireworks Finds Numerical Misalignment Can Crash RL Training Rewards(2 posts)→

Original post →

More from Infra

Infra channel →