Qwen's TRACE enables FP4 RL training of MoE models with 5.4x rollout speedup
Qwen · hf · 2026-10-07
The Qwen team proposes TRACE, an FP4 quantization framework for RL training of Mixture-of-Experts language models.
- Pain point: existing FP4 RL methods optimize train-side and rollout-side quantization independently rather than reducing the discrepancy between the two quantized execution paths.
- Method: TRACE uses rollout-guided quantization-aware training, letting rollout-side quantization outcomes guide train-side FP4 rounding decisions, plus an efficient caching scheme that selectively retains mantissa and scale information from deeper layers.
- Results: on four large-scale MoE models, joint FP4 weight/activation and FP4 KV-cache rollout matches BF16 rollout RL performance with up to 5.4x rollout speedup, and beats post-hoc FP4 quantization of BF16-trained policies.
More from Infra
- Intel CEO says Intel will stay in Musk's Terafab project after TSMC talk spooked investors — rohanpaul_ai · 2026-10-07
- Dev ships Vulkan-only local generative art app with no telemetry or cloud — ogimaru · 2026-10-07
- AMD ships ROCm 10.1, targeting storage-to-GPU data movement as the new training bottleneck — AccBalanced · 2026-10-07
- Crusoe's Path: From Stranded-Gas Bitcoin Mining to Prefab Gigawatt-Scale GPU Datacenters — AccBalanced · 2026-10-07
- Musk bets every 5GW of added US power equals roughly 1% GDP growth — XFreeze · 2026-10-07
- Burn 0.22 released: biggest Rust DL framework update yet, 6-15x faster rebuilds — JosephJacks_ · 2026-10-07