Writing a B200 attention kernel from scratch to near-SOTA, in 60 diagrams
magoghm · hn · 2026-09-02
A diagram-heavy tutorial walks through writing an attention kernel from scratch on NVIDIA's B200 to near-SOTA performance, covering TMA, tensor cores, shared memory swizzling, tiling, and warp specialization step by step.
More from Infra
- Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing — GabGarrett · 2026-09-03
- Databricks pitches agent-native data infrastructure, Lakebase Postgres at VLDB 2026 — matei_zaharia · 2026-09-03
- Lablup, Maker of GPU Orchestrator Backend.AI, Joins PyTorch Foundation as Silver Member — PyTorch · 2026-09-03
- Cursor cloud agents can now run on your own infrastructure, Mac Minis included — mattyp · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- Analyst: NVIDIA Could Become Intel Foundry's 'Customer Zero' as a Second Source Beyond TSMC — BenBajarin · 2026-09-03