Classic CUDA tutorial walks from naive matmul to 94% of cuBLAS performance step by step
Abhishekcur · x · 2026-10-01
The author recommends Si Boehm's classic CUDA optimization worklog as the resource that got him started with kernel optimization — ideal for beginners learning to write CUDA kernels.
The article uses matrix multiplication (arguably the most important algorithm on GPUs, accounting for nearly all FLOPs in LLM training and inference) to teach GPU performance fundamentals: starting from a naive kernel (309 GFLOPs/s, 1.3% of cuBLAS) and iteratively applying optimizations to reach 94% of cuBLAS FP32 performance:
- Global memory coalescing: 8.5%
- Shared memory caching: 12.8%
- 1D/2D block tiling: 36.5% → 68.7%
- Vectorized memory access: 78.4%
- Autotuning + warptiling: 93.7%
All kernel code is on GitHub, covering coalescing, shared memory, tiling, and vectorized loads.
More from Infra
- Anthropic's S-1 reveals $42B Broadcom facility backing a $125.2B TPU compute lease — roll0ver · 2026-10-01
- SpaceX's AI unit reportedly held talks to lease computing capacity to Microsoft — pstAsiatech · 2026-10-01
- HF engineer says SGLang is now his default inference framework — TheZachMueller · 2026-10-01
- Cloudflare Basin, its Iceberg-based serverless data platform, is now generally available — ritakozlov · 2026-10-01
- Cloudflare Birthday Week Day 4: edge event streaming K2, AI Search GA, Artifacts beta and more — threepointone · 2026-10-01
- Strata runs 125B Qwen3.8-Flash-Next on a 16GB consumer GPU — evilsocket · 2026-10-01