humans& Releases 4-bit RL Training Recipe

xiaosun86 · x · 2026-07-11

humans& shared their training approach: they aim to train models based on the long-term consequences of interactions between models and humans, hence their emphasis on **long-horizon multi-agent RL**. The post also mentions that they have open-sourced a **hardware-native 4-bit RL recipe** designed to significantly accelerate training.

Related event: Humans& Open-Sources 4-bit RL Training Recipe(6 posts)→

Original post →

More from Research

Research channel →