SGLang v0.5.15 Boosts Inference Throughput
TheZachMueller · x · 2026-07-15
SGLang released v0.5.15, focusing heavily on server-side optimizations for production inference:
- Production serving tuning for GLM-5.2 NVFP4, achieving 500+ tok/s/user on 8x B300 and 450 tok/s/user on 4x GB300 (bs=1).
- Added support for several new models, including Hunyuan 3 (Hy3), HRM-Text, NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, and Qwen3.6.
- Technical highlights of this update include:
- Breakable CUDA Graph is now the default capture path
- Built-in web search powered by Exa
- Added decode context parallelism for MLA models, covering DeepSeek V3
- Enhanced all-to-all communication capabilities for routed models
The original post also mentioned that run commands and a more comprehensive technical blog post will be provided later.
Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→
More from coding & agent
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11
- The browser main thread is expensive: a practical guide to JavaScript and CSS animation cost — jh3yy · 2026-09-11
- Claude Unlimited: open-source local proxy rotates accounts and API keys to keep Claude Code sessions alive — Similar_Injury_6739 · 2026-09-11
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11