ThunderKittens lands on NVIDIA Vera Rubin, pushing NVFP4 GEMMs past 22 PFLOPS
togethercompute · x · 2026-09-11
Together AI's kernels team got early access to NVIDIA's NVL72, dug through the new ISA, and brought NVFP4 and FP8 GEMMs to Vera Rubin via ThunderKittens — past 22 and 12 PFLOPS respectively, competitive with cuBLAS + CUTLASS DSL. The team notes Rubin's tensor cores consume operands twice as fast, but simply widening the MMA wasn't enough: they increased data reuse and deepened pipelines to push a 16k NVFP4 GEMM to 22.4 PFLOPS.
Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→
More from Infra
- Bartowski Details New Per-Tensor Layout Maps for GGUF Quantization in llama.cpp — bartowski1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11
- DOJ scrutinizes Nvidia's ~$20B Groq licensing deal over merger-review evasion — eyishazyer · 2026-09-11
- Persimmon Built on NVIDIA's 550B Nemotron 3 Ultra with Thousands of Blackwell GPUs — niloofar_mire · 2026-09-11
- NVIDIA details EPD disaggregation: up to 5x faster TTFT and 7x faster responses for multimodal serving — NVIDIAAI · 2026-09-11