TRIAGE stabilizes native NVFP4 RL training, hits full-precision quality at 2.3x throughput
InfiX-ai · hf · 2026-10-08
InfiX-ai's new paper tackles instability in low-precision (native NVFP4, W4A4) reinforcement learning for LLMs:
- Mismatch between learner and sampler execution destabilizes policy optimization, and mismatch magnitude alone can't identify which updates are harmful.
- The authors characterize how mismatch interacts with the policy-gradient direction, distinguishing amplifying vs. contracting updates; native NVFP4 runs show an early imbalance favoring negative-advantage, negative-gap updates whose tail tokens concentrate in a few response segments before spreading globally.
- TRIAGE performs segment-level diagnosis to selectively rebalance policy-gradient updates with bounded repair, while keeping W4A4 forward execution on both sampler and learner.
- On Qwen3-4B and Qwen3-30B-A3B it trains stably, matches full-precision performance across five math reasoning benchmarks, and delivers up to 2.3x rollout throughput over BF16.
More from Infra
- WSJ: Broadcom arranging $50B+ financing for OpenAI's custom chips under secret Nexus program — rohanpaul_ai · 2026-10-08
- Reddit asks: have AI scaling laws hit their limit, or is the compute buildout just starting? — StupidDialUp · 2026-10-08
- WSL Containers Now Generally Available: Run Linux Containers Natively on Windows — pavandavuluri · 2026-10-08
- ai& Says It's Japan's Largest Dedicated Inference Provider, Teases Post-Training Offerings — DavidBennett__ · 2026-10-08
- Microsoft unveils Surface Laptop Ultra with Nvidia RTX Spark SoC from $2,599 — Ars Technica AI · 2026-10-08
- DeepSeek V4.1 shrinks cache 437x, Flash beats V4-Pro 39 vs 36 at half the cost — DeepLearningAI · 2026-10-08