skyrl v0.4.0 runs full-context RL on a 1T-parameter model with just 16 B300 GPUs
casper_hansen_ · x · 2026-10-03
skyrl released v0.4.0, claiming full-context RL training on a 1-trillion-parameter model with just 16 B300 GPUs (2 nodes) — reportedly one of the most memory-efficient implementations available.
The release adds stable support for new models: GLM 5.3 Flash and GLM 5.2/5.3 (via trajectorylabs), Kimi K2.6/2.7 (on just 2 B300 nodes), Qwen3.8, and Nemotron 3.5 Lightning.
More from coding & agent
- Building a Self-Tagging Digital Garden So AI Agents Can Actually Use Your Bookmarks — floguo · 2026-10-03
- Karpathy reportedly says prompting is fading: graphs are the layer that survives — msharmas · 2026-10-03
- Dev: With Enough Compute, One Person Could Produce Centuries' Worth of Software — cephaloform · 2026-10-03
- OpenAI DevDay notes: 4 key levers to cut costs and boost agent performance — omarsar0 · 2026-10-03
- GitHub Copilot CLI v1.0.92-3 adds Ctrl+E picker to switch local and cloud runs — copilot-cli-release-app[bot] · 2026-10-03
- PhD-turned-founder: LLMs flipped which criteria kill programming tool startups — jimmykoppel · 2026-10-03