Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS
Together AI's kernel team gained early access to NVIDIA's Vera Rubin NVL72 system and, within days of studying the new ISA, ported ThunderKittens' NVFP4 and FP8 GEMM kernels to the Rubin GPU, achieving over 22 and 12 PFLOPS respectively—performance on par with official solutions like cuBLAS and CUTLASS. The team also published a technical blog detailing the porting process, one of the first publicly available third-party Rubin kernel optimization efforts.
Confirmed
- Together AI got early access, completed studying the new ISA and the port within days, with NVFP4 GEMM reaching 22 PFLOPS and FP8 reaching 12 PFLOPS; at 16k scale, NVFP4 GEMM was pushed up to 22.4 PFLOPS.
- The blog starts from the original B200 GEMM and progressively layers in Rubin's new features: wider MMA steps, more tensor memory/shared memory, B-side collectors, early A release, and more.
- Performance is comparable to cuBLAS and CUTLASS.
Why it matters
- The team notes Rubin's biggest change is that tensor cores consume operands twice as fast, but simply widening the MMA isn't enough to keep the cores fed; only by simultaneously improving data reuse and deepening pipelines could the 16k NVFP4 GEMM reach 22.4 PFLOPS—directly valuable reference for other kernel developers preparing to adapt to Rubin.
- The authors explicitly state these kernels are still early-stage, with further optimizations like LUT still to come, indicating Rubin's performance ceiling has yet to be fully tapped.
2026-09-11 ~ 2026-09-11 · 6 related posts
Primary sources
- ThunderKittens lands on NVIDIA Vera Rubin, pushing NVFP4 GEMMs past 22 PFLOPS — togethercompute ·
- Feeding Rubin's tensor cores took more than wider MMAs: reuse and deeper pipelines — togethercompute ·
- [source] ThunderKittens lands on NVIDIA Vera Rubin, pushing NVFP4 GEMMs past 22 PFLOPS — togethercompute · 2026-09-11
- [source] Feeding Rubin's tensor cores took more than wider MMAs: reuse and deeper pipelines — togethercompute · 2026-09-11
- Together AI's Rubin kernel blog builds from B200 GEMMs, stacking new ISA features — togethercompute · 2026-09-11
- Together Compute ships early Rubin GPU kernels, unlocking wider MMA steps and new memory features — togethercompute · 2026-09-11
2 near-duplicate retellings: simran_s_arora · HazyResearch