Harvard CS249r ML Systems Vol II performance chapter covers graph compilation, fusion, CUDA profiling

blaizedsouza · x · 2026-09-14

A GPU programming learner (Day 252/365) recommends the Performance Engineering chapter of Harvard's CS249r Machine Learning Systems Vol II, calling it remarkably thorough for a course book: graph compilation, fusion, precision, CUDA streams profiling, plus a case study and common pitfalls. The prior day's post highlighted annotated PTX/SASS for a naive matmul kernel, producer/consumer TMA pipeline diagrams, and thread block clusters reducing L2 traffic.

Original post →

More from Infra

Infra channel →