From 4 to 285 TFLOP/s: Part 1 of writing speed-of-light GEMM kernels on Blackwell B200

HanGuo97 · x · 2026-09-05

Luke D. Huang published part 1 of a blog series: writing a sequence of progressively optimized BF16 GEMM kernels in CuTeDSL on B200, ultimately reaching 99% of cuBLAS performance on several matrix dimensions. This 19-minute post covers:

The author spent the past six months deep in GPUs and performance engineering, documenting the path to speed-of-light kernels.

Original post →

More from Infra

Infra channel →