Custom CUDA SGEMM kernel matches cuBLAS, wins on 16384x16384+ matrices
goyal__pramod · x · 2026-10-02
As a feasibility study for a general tensor reduction library, the author wrote a single-precision CUDA matmul kernel with bit-for-bit identical results to cuBLAS. It's a few percent faster on square matrices of 16384x16384 and larger, slightly slower on smaller ones, with performance varying by power profile, clocks, and compiler. Two notable techniques: scheduling thread blocks on a Hilbert curve to maximize L2 cache efficiency, and warp-tile arrangements that sync only across half then quarter of the thread block instead of the whole block. The post covers setup, implementation overview, benchmarking, and full code, with detailed kernel teardowns to follow.
Related event: Hand-Written CUDA Matmul Kernels Approach and Beat cuBLAS(3 posts)→
More from Infra
- AI Capex Now Exceeds the Railroad Boom's GDP Share, Yet Demand Lags — FinanceYF5 · 2026-10-02
- Real-time AI serving costs up to 56x more than needed — one engine hits 56 sessions per H100 — Ok_boss_labrunz · 2026-10-02
- Cloudflare launches SQL API to query Workers logs and traces, replacing GraphQL plans — irvinebroque · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- Lightpanda 1.0 ships: a Zig-built browser for AI agents, out of beta — jedisct1 · 2026-10-02
- Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s — W61k3r · 2026-10-02