SGLang v0.5.15 Released
BanghuaZ · x · 2026-07-11
SGLang has released v0.5.15, with a major focus on inference serving optimizations for production environments.
Key Updates
- The team optimized production serving around GLM-5.2 NVFP4, achieving 500+ tok/s/user on 8x B300 and 450 tok/s/user on 4x GB300 (bs=1).
- Newly supported models include: Hunyuan 3, HRM-Text, NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, Qwen3.6, etc.
- The default capture path has been changed to Breakable CUDA Graph.
- Built-in web search is now available, powered by Exa.
- Enhanced decode context parallelism, covering MLA models (including DeepSeek V3).
- Added FlashInfer all-to-all for routed MoE.
- DeepSeek-V4 related FlashMLA sparse prefill is now enabled by default, improving long-context prefill and end-to-end performance.
Additional Info
- This version welcomes 43 new contributors.
Related event: SGLang v0.5.15 Boosts Production Inference Performance(2 posts)→
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11