From 4 to 285 TFLOP/s: Part 1 of writing speed-of-light GEMM kernels on Blackwell B200
HanGuo97 · x · 2026-09-05
Luke D. Huang published part 1 of a blog series: writing a sequence of progressively optimized BF16 GEMM kernels in CuTeDSL on B200, ultimately reaching 99% of cuBLAS performance on several matrix dimensions. This 19-minute post covers:
- Kernel 1 (naive matmul): one thread per output element walking the full K dimension, a 4 TFLOP/s baseline.
- A CuTeDSL mental model: how it describes data layout, tiles data, and maps tiles onto GPU operations.
- Kernel 2 (tensor cores, TMA, TMEM): per-CTA tiles of C, TMA with barriers moving A/B tiles into shared memory, tcgen05 MMA accumulating into TMEM, and a per-thread epilogue through registers — 155 TFLOP/s.
- Kernel 3 (swizzling): identifying shared memory bank conflicts as the bottleneck and changing the layout only — 285 TFLOP/s.
The author spent the past six months deep in GPUs and performance engineering, documenting the path to speed-of-light kernels.
More from Infra
- Free tokens are fueling open-source and local AI, Jason argues citing Jensen — AccBalanced · 2026-09-05
- GPU Sandboxes as the Compute Primitive for Recursive Self-Improvement — AAAzzam · 2026-09-05
- SemiAnalysis and vLLM launch AgentX, a benchmark for real multi-turn agentic traffic — AccBalanced · 2026-09-05
- Zilliz CTO on Notion's Vector Infra: 90% Cost Cuts and the Platform Hiding Under One AI Feature — J_Luan_ · 2026-09-05
- 36B MoE open model K2-Horizon with 4B active params ships as local-runnable GGUF — rupspace · 2026-09-05
- The longer agent memory lives, the riskier a single hot retrieval index becomes — Cautious_Bit_8521 · 2026-09-05