GPU Programming Diary: Revisiting the Classic CUDA Matmul Optimization Worklog and MIT's Sparsity Lecture

NandoDF · x · 2026-09-20

A learner documenting a year of GPU programming highlights Simon Boehm's 2022 worklog on optimizing a CUDA matmul kernel to cuBLAS-like performance — one of the most-recommended resources outside PMPP and GPU Mode — valued less for cutting-edge techniques than for watching a great performance engineer think step by step. Day 258 covered MIT's Efficient AI Computing Lecture 3 on hardware sparsity support, pruning granularity, N:M sparsity, and scaling factors, with assignments available in Colab.

Original post →

More from Infra

Infra channel →