Writing a B200 attention kernel from scratch to near-SOTA, in 60 diagrams

magoghm · hn · 2026-09-02

A diagram-heavy tutorial walks through writing an attention kernel from scratch on NVIDIA's B200 to near-SOTA performance, covering TMA, tensor cores, shared memory swizzling, tiling, and warp specialization step by step.

Original post →

More from Infra

Infra channel →