SGLang v0.5.15 Released
ying11231 · x · 2026-07-15
The quoted post highlights the release of SGLang v0.5.15, focusing on optimizations for production-grade inference services.
This update includes:
- Production tuning for GLM-5.2 NVFP4, achieving 500+ tok/s/user on 8x B300 and 450 tok/s/user on 4x GB300 (bs=1)
- Plans to share run commands and full technical details in upcoming threads and blogs
- Added support for multiple new models: Hunyuan 3, HRM-Text, NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, Qwen3.6
- Other major updates: Breakable CUDA Graph default capture path, built-in web search (powered by Exa), decode context parallelism for MLA models, and FlashInfer all-to-all
Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→
More from Infra
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21
- llama.garden is using torrents and web seeds to decentralize LLM distribution — de4dee · 2026-07-21
- One command finds which of hundreds of models fit your hardware — AlexsJones · 2026-07-21