SkyRL v0.4 trains 1T-param Kimi K2.7 with RL on just 16 B300 GPUs
casper_hansen_ · x · 2026-10-03
SkyRL released v0.4.0 of its open-source RL post-training framework, focused on large-scale efficiency.
- Ultra-low hardware threshold: Via LoRA + INT4 inference serving + fake-quantization on the trainer to reduce train/inference mismatch, the 1T+ parameter Kimi K2.6/2.7 can now run single-turn RL on just 2 B300 nodes (16 GPUs) — the author notes most frameworks demand 16 full nodes and may not even offer full context.
- New model support: GLM 5.3 Flash and GLM 5.3 (in collaboration with Trajectory), Kimi K2.6/2.7, Qwen3.8, Nemotron 3.5-Lightning.
- RL stability: Top-P sampler replay plus optimized rollout router replay (R3) data transfer for large MoE models, contributed by Dominic Yurk of Lila Sciences.
- Weight sync and throughput: Full FP8/MXFP8 RL, Delta Weight Sync, sharded weight transfer via Ray Direct Transfer + NIXL, plus large-scale SFT and Tinker Server scalability improvements.
More from coding & agent
- Browser sidebar built on opencode v2 lets AI run scripts and control any page — uwukko · 2026-10-03
- 'next-steps' Claude Mod Tames the Chaos of Running 10 Claudes in Parallel — daniel_mac8 · 2026-10-03
- Dot agent proactively flags CI failures from GitHub's macos-14 runner retirement — charliermarsh · 2026-10-03
- KNOWS: computer-use agents that research the web and build real Google Docs, Sheets and Slides — anmarasovic · 2026-10-03
- AWS open-sources Strands Decider 2B, a 2B decision model for fast agentic routing — amaarora · 2026-10-03
- Engineer uses Claude Opus as an adversarial reviewer to settle technical disputes — keyanzhang · 2026-10-03