Humans& Open-Sources 4-bit RL Training Recipe
The Humans& team has released an open-source, hardware-native 4-bit Reinforcement Learning (RL) training recipe. The core goal of this recipe is to optimize long-horizon, multi-agent RL, enabling models to learn directly from the feedback and consequences of long-term interactions with humans.
Key Details and Technical Advantages
The approach emphasizes maintaining quantization consistency between the training and rollout phases, utilizing dequantized backward propagation. According to the team, this NVFP4 approach does not suffer from performance degradation compared to 8-bit methods, while significantly accelerating the training process. Currently, the training recipe has been implemented and is supported by the Miles / SGLang ecosystem. Additionally, alongside the open-source release, the team shared detailed blogs and visualizations on quantization to make the topic more accessible.
2026-07-11 ~ 2026-07-11 · 6 related posts
- humans& Releases 4-bit RL Training Recipe — xiaosun86 · 2026-07-11
- Open-Source NVFP4 Reinforcement Learning Solution — gharik · 2026-07-11
- Open-Source 4-bit RL Training Recipe — gharik · 2026-07-11
- Open-Sourced 4-bit RL Training Recipe — hsu_byron · 2026-07-11
- Open-Sourced 4-bit RL Recipe Boosts Training Efficiency — niloofar_mire · 2026-07-11
1 near-duplicate retellings: hyunw_kim