SkyRL v0.4 Enables RL Training of Trillion-Parameter Models on Just 16 B300 GPUs
Open-source RL post-training framework SkyRL released v0.4.0, claiming that with LoRA, INT4 inference serving and pseudo-quantization, full-context RL training of a 1-trillion-parameter model (Kimi K2.7) is possible on just 16 B300 GPUs across 2 nodes, one of the most memory-efficient implementations to date.
2026-10-03 ~ 2026-10-03 · 2 related posts
- skyrl v0.4.0 runs full-context RL on a 1T-parameter model with just 16 B300 GPUs — casper_hansen_ · 2026-10-03
- SkyRL v0.4 trains 1T-param Kimi K2.7 with RL on just 16 B300 GPUs — casper_hansen_ · 2026-10-03