skyrl v0.4.0 runs full-context RL on a 1T-parameter model with just 16 B300 GPUs

casper_hansen_ · x · 2026-10-03

skyrl released v0.4.0, claiming full-context RL training on a 1-trillion-parameter model with just 16 B300 GPUs (2 nodes) — reportedly one of the most memory-efficient implementations available.

The release adds stable support for new models: GLM 5.3 Flash and GLM 5.2/5.3 (via trajectorylabs), Kimi K2.6/2.7 (on just 2 B300 nodes), Qwen3.8, and Nemotron 3.5 Lightning.

Related event: SkyRL v0.4 Enables RL Training of Trillion-Parameter Models on Just 16 B300 GPUs(2 posts)→

Original post →

More from coding & agent

coding & agent channel →