Classic CUDA tutorial walks from naive matmul to 94% of cuBLAS performance step by step

Abhishekcur · x · 2026-10-01

The author recommends Si Boehm's classic CUDA optimization worklog as the resource that got him started with kernel optimization — ideal for beginners learning to write CUDA kernels.

The article uses matrix multiplication (arguably the most important algorithm on GPUs, accounting for nearly all FLOPs in LLM training and inference) to teach GPU performance fundamentals: starting from a naive kernel (309 GFLOPs/s, 1.3% of cuBLAS) and iteratively applying optimizations to reach 94% of cuBLAS FP32 performance:

All kernel code is on GitHub, covering coalescing, shared memory, tiling, and vectorized loads.

Original post →

More from Infra

Infra channel →