Hand-Written CUDA Matmul Kernels Approach and Beat cuBLAS

A classic CUDA matmul optimization tutorial trending in the community reaches 94% of cuBLAS performance, while a hand-written SGEMM kernel matches cuBLAS bit-for-bit and beats it on 16384x16384 matrices.

2026-10-01 ~ 2026-10-02 · 3 related posts