SGLang lands on NVIDIA Vera Rubin, speeds up Kimi K3 inference by up to 20%
charles_irl · x · 2026-10-10
The SGLang project, working with NVIDIA, ported the inference engine to early-access Vera Rubin hardware and optimized attention, MoE, and speculative verification kernels for Kimi K3: up to 20% faster FP8 MLA at batch 1 / 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. SGLang also powers RL rollouts, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Related event: SGLang Ported to NVIDIA Vera Rubin, Speeding Kimi K3 Inference Up to 20%(2 posts)→
More from Infra
- Alchemy + Cloudflare combo enables blazing-fast iteration, no local dev server needed — samgoodwin89 · 2026-10-10
- FT: OpenAI's $70B ARR is investor-grossed; accounting choices hide billions in gap — rohanpaul_ai · 2026-10-10
- Infra engineers joke about outlawing all dynamism in frontier models — charles_irl · 2026-10-10
- Basalt: open-source Blackwell inference engine hits 665 tok/s, 2.6x Strata on a 5090 rig — jesdga95 · 2026-10-10
- Supabase outpaced Vercel in site visits for 12 straight months as AI builder layer doubles — FinanceYF5 · 2026-10-10
- Google AI Pro plans bundle Colab A100 80GB access, enabling full fine-tunes of Gemma up to 31B — algo_diver · 2026-10-10