SGLang on NVIDIA Vera Rubin: up to 20% faster Kimi K3 inference, 4.8x gains for Cognition

ying11231 · x · 2026-10-10

The LMSYS team got early access to two NVIDIA Vera Rubin nodes (8 GPUs) and ported SGLang (inference) and Miles (RL training) to the next-gen platform, optimizing kernels for Kimi K3 NVFP4 — a 2.8T-parameter model with 1M context, 69 KDA + 24 MLA layers, and an 896-expert LatentMoE:

Vera Rubin's key changes vs Hopper/Blackwell: 327 KiB shared memory per CTA (up from 227 KiB), 212 SMs, and NVLink 6. Cognition has already deployed Vera Rubin NVL72 with its own SGLang-based stack, reporting up to a 4.8x total token throughput increase over GB200 NVL72.

Original post →

More from Infra

Infra channel →