SkyRL v0.4 Enables RL Training of Trillion-Parameter Models on Just 16 B300 GPUs

Open-source RL post-training framework SkyRL released v0.4.0, claiming that with LoRA, INT4 inference serving and pseudo-quantization, full-context RL training of a 1-trillion-parameter model (Kimi K2.7) is possible on just 16 B300 GPUs across 2 nodes, one of the most memory-efficient implementations to date.

2026-10-03 ~ 2026-10-03 · 2 related posts