SGLang lands on NVIDIA Vera Rubin, speeds up Kimi K3 inference by up to 20%

charles_irl · x · 2026-10-10

The SGLang project, working with NVIDIA, ported the inference engine to early-access Vera Rubin hardware and optimized attention, MoE, and speculative verification kernels for Kimi K3: up to 20% faster FP8 MLA at batch 1 / 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. SGLang also powers RL rollouts, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

Related event: SGLang Ported to NVIDIA Vera Rubin, Speeding Kimi K3 Inference Up to 20%(2 posts)→

Original post →

More from Infra

Infra channel →