SGLang v0.5.15 Released
BanghuaZ · x · 2026-07-11
SGLang has released v0.5.15, with a major focus on inference serving optimizations for production environments.
Key Updates
- The team optimized production serving around GLM-5.2 NVFP4, achieving 500+ tok/s/user on 8x B300 and 450 tok/s/user on 4x GB300 (bs=1).
- Newly supported models include: Hunyuan 3, HRM-Text, NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, Qwen3.6, etc.
- The default capture path has been changed to Breakable CUDA Graph.
- Built-in web search is now available, powered by Exa.
- Enhanced decode context parallelism, covering MLA models (including DeepSeek V3).
- Added FlashInfer all-to-all for routed MoE.
- DeepSeek-V4 related FlashMLA sparse prefill is now enabled by default, improving long-context prefill and end-to-end performance.
Additional Info
- This version welcomes 43 new contributors.
Related event: SGLang v0.5.15 Boosts Production Inference Performance(2 posts)→
More from Infra
- Vercel AI Gateway data shows Anthropic, OpenAI and Google at 97.09% spend share — cramforce · 2026-07-21
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Microsoft and Mistral sign multi-billion-dollar deal to expand AI infrastructure in Europe — The Decoder · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21