Unified FP8 in Training and Rollout Speeds Up RL by 16%
joecole · x · 2026-07-30
A new paper Jet-RL from NVIDIA, MIT, and other institutions investigates the stability of using FP8 quantization during Reinforcement Learning (RL) training for Large Language Models (LLMs).
- The Bottleneck: The rollout phase typically consumes over 70% of total RL training time. The common practice of using BF16 for training and FP8 for rollout introduces significant numerical mismatch, leading to severe instability and accuracy collapse.
- The Jet-RL Solution: The framework proposes a unified FP8 precision flow for both training and rollout, minimizing numerical discrepancies and avoiding inefficient inter-step calibration.
- Performance: Experiments show that this approach achieves up to a 33% speedup in the rollout phase, a 41% speedup in the training phase, and a 16% end-to-end speedup over BF16 training while maintaining robust accuracy.
More from Infra
- SK Hynix Earnings Analysis: AI Memory Demand Strong, Market Overreacts to Oversupply — tengyanAI · 2026-07-30
- Single 8x 5090 Rig Hits 167k tokens/s Training Throughput, Beating DDP — jon_durbin · 2026-07-30
- Buildcleaner reclaims 443GB of disk space by cleaning build artifacts, free and open-source MIT — jasonkneen · 2026-07-30
- LLM Inference Costs Drop Below $3 with B200s, Yet API Prices Stay High — AccBalanced · 2026-07-30
- Bought GPUs to Escape API Fees, Realized a Single RTX 5090 Is Enough — Ok-Shower7286 · 2026-07-30
- Yann LeCun and Others Discuss: LLMs are the New Compilers, Performance is a Function of Compute — yisongyue · 2026-07-30