SGLang v0.5.18: Engine 2.38x Faster, AMD NVFP4 Support Added

BanghuaZ · x · 2026-08-25

SGLang v0.5.18 is released with 710 PRs from 212 contributors. Key highlights include engine startup speedup of 2.38x via disk streaming and CUDA graph parallelization. Support for new models like Muse Glimmer and Intern-S2-Mobius is added. All compiled kernels are now unified in one directory. NVFP4 checkpoints can now run on AMD GPUs via online MXFP4 requantization, maintaining 97.5-100% accuracy for models like DeepSeek-R1. Additionally, Kimi K3 sees 1.37-1.77x throughput gains on AMD MI355X, and DeepSeek-V4 decoding is optimized with FlashInfer.

Original post →

More from Infra

Infra channel →