GPU Programming Diary: Revisiting the Classic CUDA Matmul Optimization Worklog and MIT's Sparsity Lecture
NandoDF · x · 2026-09-20
A learner documenting a year of GPU programming highlights Simon Boehm's 2022 worklog on optimizing a CUDA matmul kernel to cuBLAS-like performance — one of the most-recommended resources outside PMPP and GPU Mode — valued less for cutting-edge techniques than for watching a great performance engineer think step by step. Day 258 covered MIT's Efficient AI Computing Lecture 3 on hardware sparsity support, pruning granularity, N:M sparsity, and scaling factors, with assignments available in Colab.
More from Infra
- Data center tax debate lacks a baseline: Tax Foundation data on billion-dollar facility rates — AndyMasley · 2026-09-20
- Are data centers dodging taxes? Tax Foundation data on $1B facilities says no — AndyMasley · 2026-09-20
- AMD carries out serious software optimizations for Kimi-K3, analyst says — AccBalanced · 2026-09-20
- RAM price jumps from $350 to $640 in months, pricing out new PCs — chrisalbon · 2026-09-20
- NVIDIA engineer breaks down why DeepSeek re-engineered V4.1 Flash for speed — thursdai_pod · 2026-09-20
- $140 Radeon MI50 paired with GTX-1080Ti boosts local 27B-35B LLM speeds up to 9x — tabletuser_blogspot · 2026-09-20