Harvard CS249r ML Systems Vol II performance chapter covers graph compilation, fusion, CUDA profiling
blaizedsouza · x · 2026-09-14
A GPU programming learner (Day 252/365) recommends the Performance Engineering chapter of Harvard's CS249r Machine Learning Systems Vol II, calling it remarkably thorough for a course book: graph compilation, fusion, precision, CUDA streams profiling, plus a case study and common pitfalls. The prior day's post highlighted annotated PTX/SASS for a naive matmul kernel, producer/consumer TMA pipeline diagrams, and thread block clusters reducing L2 traffic.
More from Infra
- No.2 US law firm Latham & Watkins buys Nvidia servers to fine-tune open models in-house — ai · 2026-09-14
- Andrew Chen's homelab for local AI: 5090 eGPU, dual DGX Spark and a routing plugin — andrewchen · 2026-09-14
- Second-largest US law firm may spend hundreds of millions fine-tuning open-weight Nemotron on its own GPUs — ivan_bezdomny · 2026-09-14
- Palantir names Nebius its preferred sovereign AI infrastructure partner — pdamodaran · 2026-09-14
- INT21's SwarmOS runs 10,000 agents to evolve and optimize inference engines — bingxu_ · 2026-09-14
- Flawed Routers Flood University of Wisconsin Internet Time Server (2003) — walrus01 · 2026-09-14