ThunderKittens now runs on NVIDIA Rubin: 22 and 12 PFLOPS on nvfp4/fp8 GEMMs, rivaling cuBLAS
simran_s_arora · x · 2026-09-11
Together AI's kernels team got early access to NVIDIA Vera Rubin NVL72 and ported ThunderKittens to Rubin in days, digging through the new ISA to bring up nvfp4 and fp8 GEMMs—pushing past 22 and 12 PFLOPS respectively, competitive with cuBLAS and the Cute DSL.
Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→
More from Infra
- L3Harris says fine-tuned open-source models beat frontier AI in 48 hours at 95% lower cost — eliano · 2026-09-11
- Qwen3-TTS 1.7B hits 1.6x real-time voice cloning on CPU via llama.cpp — alexcovo_eth · 2026-09-11
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Qwen 3.8 125B-A6B runs 15.3% faster on Mac via speculative decoding on mlx.fast — TheMoonMidas · 2026-09-11
- Olam Labs CEO: only compute and data remain as bottlenecks to AGI — garrytan · 2026-09-11
- Open-sourced inference acceleration for structure-based models ships benchmarked and documented — AllThingsApx · 2026-09-11