SGLang Team Open-Sources Miles v0.1: 744B Model RL on 64 GPUs at 263s/Step
aigclink · x · 2026-09-13
The team behind SGLang has open-sourced Miles v0.1, a production-ready post-training framework that turns training a company-specific model from a costly gamble into an engineering task — spin it up with Docker and swap in your own data and tasks.
Key capabilities
- Focused on agentic RL: models learn by trial and error in real environments, using tools, operating terminals, and completing multi-step tasks — aiming for a "worker" rather than a "chatbot."
- Scales to frontier models: the paper's case study runs GLM-5.2 (744B-A40B) on terminal-use coding tasks across 64 NVIDIA GB300 GPUs, with a median step time of 263 seconds over the first 30 steps.
- Reliability: engineered against mid-run crashes that waste huge compute; Modal calls it "consistently stable over long runs."
- One framework covers SFT, RL, distillation, and LoRA, and extends to diffusion models.
- Data stays local: weights and training logs remain on-premises, a must for finance, healthcare, and government clients.
- Supports DeepSeek, Qwen, GLM, Kimi, Gemma, and Nemotron, on both NVIDIA and AMD GPUs.
- Already used in production by Periodic Labs (trillion-parameter training on thousands of GPUs), Modal, IBM, Decagon, and Nebius.
The author argues startups should seriously consider vertical models, enterprises now have controllable private-training costs, and AI service providers/FDEs may find a new market.
Related event: SGLang Team Open-Sources Miles v0.1 Production Post-Training Framework(2 posts)→
More from Infra
- Why someone thinks Huawei should build a desktop inference box with 1TB of VRAM — AIFlow_ML · 2026-09-13
- 5+ specialized inference engines shipped in a month, sparking ecosystem fragmentation debate — JustinLin610 · 2026-09-13
- Cloudflare CEO: One agent per knowledge worker would need 40x the world's CPUs — BenBajarin · 2026-09-13
- GACS2026 AI Chip Summit Sets Shanghai Agenda with Agent and Embodied-Intelligence Tracks — 智东西 · 2026-09-13
- Local LLM optimists vindicated: early GPU buyers wish they'd bought more — nptacek · 2026-09-13
- Reading up on terahertz pulse reflectometry for inspecting CoWoS packaging layers — jwt0625 · 2026-09-13