Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS

Together AI's kernel team gained early access to NVIDIA's Vera Rubin NVL72 system and, within days of studying the new ISA, ported ThunderKittens' NVFP4 and FP8 GEMM kernels to the Rubin GPU, achieving over 22 and 12 PFLOPS respectively—performance on par with official solutions like cuBLAS and CUTLASS. The team also published a technical blog detailing the porting process, one of the first publicly available third-party Rubin kernel optimization efforts.

Confirmed

Why it matters

2026-09-11 ~ 2026-09-11 · 6 related posts

Primary sources

2 near-duplicate retellings: simran_s_arora · HazyResearch