Comprehensive Triton GPU programming lecture: from H100 internals to FlashAttention
kalyan_kpl · x · 2026-10-05
A comprehensive video lecture on GPU programming covering:
- How GPUs sit in computers: PCIe, NVLink, servers and clusters
- Inside the NVIDIA H100: SMs, FP32 lanes, warp scheduler, tensor cores, shared memory
- Threads/warps/blocks/grids and per-thread data mapping
- Memory: SRAM vs DRAM, DDR vs GDDR vs HBM, caches, latency vs bandwidth
- Optimizations: latency hiding, coalescing, kernel fusion
- Why Triton exists: thinking in tiles instead of threads
- Memory-bound vs compute-bound, arithmetic intensity, roofline model
- The N×N attention problem and how FlashAttention fixes it
More from Infra
- InP choke pushes industry toward 1um lasers on GaAs and hollow-core fiber — jwt0625 · 2026-10-05
- Cloudflare's new Web Search API is just a wrapper around Exa and a few other providers — gaganghotra_ · 2026-10-05
- Top 10 providers on OpenRouter ranked by monthly token volume — stuffyokodraws · 2026-10-05
- Crusoe CEO explains its layered compute business: 5-year leases to high-margin inference — AccBalanced · 2026-10-05
- AMD MI355X beats Nvidia on inference margins in SemiAnalysis InferenceX benchmark — AccBalanced · 2026-10-05
- Best or worst time to start an inference company? Scale economics say brutal — AccBalanced · 2026-10-05