Open-Source NVFP4 Reinforcement Learning Solution
gharik · x · 2026-07-11
This post introduces an open-source, hardware-native NVFP4 reinforcement learning training solution, emphasizing that it is backed by the Miles / SGLang ecosystem and aims to maintain quantization consistency between training and rollout.
Key details include:
- Utilizes dequantized backward
- Features 4/6 adaptive scaling with zero extra latency during rollout
- Selective precision remains consistent from checkpoint to rollout
- In experiments, the reward curve matches BF16
- Unlocks up to roughly 9x tensor core FLOPs on next-gen NVIDIA GPUs
The original poster summarizes this as the "4-bitter lesson": the true difficulty lies in the interactions between system components rather than the individual components themselves.
Related event: Humans& Open-Sources 4-bit RL Training Recipe(6 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22