ThunderKittens lands on NVIDIA Vera Rubin, pushing NVFP4 GEMMs past 22 PFLOPS

togethercompute · x · 2026-09-11

Together AI's kernels team got early access to NVIDIA's NVL72, dug through the new ISA, and brought NVFP4 and FP8 GEMMs to Vera Rubin via ThunderKittens — past 22 and 12 PFLOPS respectively, competitive with cuBLAS + CUTLASS DSL. The team notes Rubin's tensor cores consume operands twice as fast, but simply widening the MMA wasn't enough: they increased data reuse and deepened pipelines to push a 16k NVFP4 GEMM to 22.4 PFLOPS.

Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→

Original post →

More from Infra

Infra channel →