OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale

teortaxesTex · x · 2026-10-09

slime v0.4.0 cites OpenAI's Navier-Stokes effort—10,000 concurrent agents and 130B output tokens—as motivation for scaling RL infrastructure. Four architectural changes: sync training for algorithmic exploration plus fully async training for throughput; distributed rollout orchestration across cluster CPUs; a new storage middleware called straw with durable queues and shared tensor storage; and independent recovery that restarts Megatron while healthy SGLang engines keep running, preserving accepted rollout work.

Original post →

More from coding & agent

coding & agent channel →