humans& Releases 4-bit RL Training Recipe
xiaosun86 · x · 2026-07-11
humans& shared their training approach: they aim to train models based on the long-term consequences of interactions between models and humans, hence their emphasis on **long-horizon multi-agent RL**. The post also mentions that they have open-sourced a **hardware-native 4-bit RL recipe** designed to significantly accelerate training.
Related event: Humans& Open-Sources 4-bit RL Training Recipe(6 posts)→
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21