SGLang on NVIDIA Vera Rubin: up to 20% faster Kimi K3 inference, 4.8x gains for Cognition
ying11231 · x · 2026-10-10
The LMSYS team got early access to two NVIDIA Vera Rubin nodes (8 GPUs) and ported SGLang (inference) and Miles (RL training) to the next-gen platform, optimizing kernels for Kimi K3 NVFP4 — a 2.8T-parameter model with 1M context, 69 KDA + 24 MLA layers, and an 896-expert LatentMoE:
- Up to 20% faster FP8 MLA at batch 1 / 128K context
- 20% faster KDA verification with bitwise-identical output
- 5.9% end-to-end speedup from MoE tail fusion, removing 276 kernel launches per decode step
- Miles runs end-to-end RL out of the box, including agentic RL with 64 concurrent sandboxes on the Vera CPU
Vera Rubin's key changes vs Hopper/Blackwell: 327 KiB shared memory per CTA (up from 227 KiB), 212 SMs, and NVLink 6. Cognition has already deployed Vera Rubin NVL72 with its own SGLang-based stack, reporting up to a 4.8x total token throughput increase over GB200 NVL72.
More from Infra
- SGLang lands on NVIDIA Vera Rubin, speeds up Kimi K3 inference by up to 20% — charles_irl · 2026-10-10
- Paper argues nested von Neumann architecture can make a million processors act as one computer — bronzeagepapi · 2026-10-10
- 'CUDA cores' are a marketing scam since Kepler: SM count is the only metric that matters — blelbach · 2026-10-10
- Wasted Highway Cloverleaf Land Could Host Data Centers, Speculates Viral Thread — gregmushen · 2026-10-10
- Industrially, What You Do More Of Gets Cheaper: The Learning Curve Behind Data Center Costs — Afinetheorem · 2026-10-10
- Agent-built system hits 2,242 tok/s on AMD MI300As, 2.33× faster than SGLang in 105 hours — bariskasikci · 2026-10-10