Custom CUDA SGEMM kernel matches cuBLAS, wins on 16384x16384+ matrices

goyal__pramod · x · 2026-10-02

As a feasibility study for a general tensor reduction library, the author wrote a single-precision CUDA matmul kernel with bit-for-bit identical results to cuBLAS. It's a few percent faster on square matrices of 16384x16384 and larger, slightly slower on smaller ones, with performance varying by power profile, clocks, and compiler. Two notable techniques: scheduling thread blocks on a Hilbert curve to maximize L2 cache efficiency, and warp-tile arrangements that sync only across half then quarter of the thread block instead of the whole block. The post covers setup, implementation overview, benchmarking, and full code, with detailed kernel teardowns to follow.

Related event: Hand-Written CUDA Matmul Kernels Approach and Beat cuBLAS(3 posts)→

Original post →

More from Infra

Infra channel →