SGLang v0.5.18: Engine 2.38x Faster, AMD NVFP4 Support Added
BanghuaZ · x · 2026-08-25
SGLang v0.5.18 is released with 710 PRs from 212 contributors. Key highlights include engine startup speedup of 2.38x via disk streaming and CUDA graph parallelization. Support for new models like Muse Glimmer and Intern-S2-Mobius is added. All compiled kernels are now unified in one directory. NVFP4 checkpoints can now run on AMD GPUs via online MXFP4 requantization, maintaining 97.5-100% accuracy for models like DeepSeek-R1. Additionally, Kimi K3 sees 1.37-1.77x throughput gains on AMD MI355X, and DeepSeek-V4 decoding is optimized with FlashInfer.
More from Infra
- AMD showcases memory fabric design optimizing memory resources — BenBajarin · 2026-08-25
- AMD details MI455X GPU and Helios rack-scale system at Hot Chips — firstadopter · 2026-08-25
- Nvidia's Cost Advantage Remains Even if Competitor Chips Were Free — sinclairx · 2026-08-25
- Lium: A Vast/Runpod Competitor Offering Cheaper GPUs and No KYC — markjeffrey · 2026-08-25
- Vera Rubin cooling loop operates with just 10°C delta — beffjezos · 2026-08-25
- NVIDIA Rubin's power smoothing tech changes data center economics — BenBajarin · 2026-08-25