ThunderKittens ports to NVIDIA Vera Rubin, hitting 22 PFLOPS on NVFP4 GEMMs
HazyResearch · x · 2026-09-11
Together AI's kernels team got early access to NVIDIA's Vera Rubin NVL72 and ported ThunderKittens to the new platform in a few days.
- Reworked kernels implement nvfp4 and fp8 GEMMs on Rubin's new ISA
- Benchmarks exceed 22 and 12 PFLOPS respectively, competitive with cuBLAS + CUTLASS DSL
- The team describes the porting as smooth and published a blog with details
Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→
More from Infra
- L3Harris says fine-tuned open-source models beat frontier AI in 48 hours at 95% lower cost — eliano · 2026-09-11
- Qwen3-TTS 1.7B hits 1.6x real-time voice cloning on CPU via llama.cpp — alexcovo_eth · 2026-09-11
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Qwen 3.8 125B-A6B runs 15.3% faster on Mac via speculative decoding on mlx.fast — TheMoonMidas · 2026-09-11
- Olam Labs CEO: only compute and data remain as bottlenecks to AGI — garrytan · 2026-09-11
- Open-sourced inference acceleration for structure-based models ships benchmarked and documented — AllThingsApx · 2026-09-11