SGLang 0.5.15 Boosts Inference Speed
BanghuaZ · x · 2026-07-12
SGLang v0.5.15 has been released, focusing on performance optimizations for production inference. - Official benchmarks indicate that on **8x B300**, production serving for **GLM-5.2 NVFP4** achieves **500+ tok/s/user**; on **4x GB300**, it reaches **450 tok/s/user** (`bs=1`). - This release cycle primarily focused on inference stack tuning, with full technical details and run commands coming soon. - Newly supported models include: **Hunyuan 3 (Hy3)**, **HRM-Text**, **NVIDIA LocateAnything-3B**, **Baidu Unlimited-OCR**, **JoyEcho**, and **Qwen3.6**. - Version highlights also feature: - **Breakable CUDA Graph** becoming the default capture path - Built-in **web search** powered by **Exa** - **decode context parallelism** for **MLA** models, including **DeepSeek V3** - **FlashInfer all-to-all** for routing models
Related event: SGLang v0.5.15 Boosts Production Inference Performance(2 posts)→
More from Infra
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21
- Voice-agent teams should use platforms first, then own STT events when failures get weird — FollowingSuitable941 · 2026-07-21