GPU Mode Releases 100+ Hands-On Lectures on CUDA and Kernel Optimization
blaizedsouza · x · 2026-08-21
The gpu-mode/lectures GitHub repo aggregates over 100 hands-on lectures on GPU programming. Topics cover CUDA, Triton, FlashAttention, vLLM, NCCL, SGLang, and kernel optimization, diving into memory architecture, Tensor Cores, distributed communication, and high-performance LLM inference. Speakers are from Stanford, Princeton, Meta, NVIDIA, AMD, Intel, and the PyTorch team.
More from Infra
- The Math: Claiming 100T Tokens/Day Would Need ~580K GPUs — teortaxesTex · 2026-08-21
- Moore's Law Fading: Non-Silicon Computing and Novel Architectures to See Capital Influx — MikePFrank · 2026-08-21
- "Why do we need more datacenters? Just write faster kernels" — basedjensen · 2026-08-21
- Cornell Nested Architecture Cuts Training Compute by 36% — burkov · 2026-08-21
- Tencent releases FlashPrefill V2 for efficient long-context LLM serving — tencent · 2026-08-21
- SGLang author asks community for pain points, vows to fix them — BanghuaZ · 2026-08-21