SGLang v0.5.17 Released: Adds Support for Kimi K3 and MiniMax-H3 Video Generation
BanghuaZ · x · 2026-08-11
SGLang has released version v0.5.17, bringing several major updates:
- New Model Support: Integrated Kimi K3 (optimized for 1M context with KDA-aware prefix caching), runnable on both NVIDIA and AMD GPUs. Added MiniMax-H3, an open video-and-audio generation model capable of producing 4-15s clips at 24fps locally on 2x RTX 5090.
- Performance: MoE prefill is up to 1.92x faster, Decode Context Parallelism (DCP) got pluggable comm backends, and multi-turn agents received a session-aware KV cache.
- Engineering: The frontend server got an initial Rust rewrite, and engine restarts are faster with a weight cache.
- The release also welcomed 45 new contributors.
More from Infra
- Inside Xanadu's Lab: Ultra-low Loss Thin Film Lithium Niobate Switch Wafers — ceciletamura · 2026-08-11
- Beyond GPUs: Rethinking the Energy and Architecture Stack for Next-Gen AI Inference — prateekj · 2026-08-11
- Running Local LLMs on Strix Halo: Are 64GB/128GB RAM Variants Practical? — riklaunim · 2026-08-11
- Open Models Matching Cloud? It's Now an Engineering Tradeoff — cocktailpeanut · 2026-08-11
- GitHub Actions Outage Last Week: Users Await Incident Report — SkyLi0n · 2026-08-11
- Wall Street Giants Partner with Nvidia on $500B AI Infrastructure Financing — firstadopter · 2026-08-11