How Tri Dao's FlashAttention became a cornerstone of modern LLM training
thisdudelikesAI · x · 2026-10-01
- A 2022 paper by then-Stanford PhD student Tri Dao, written with Chris Ré and Stefano Ermon, made AI training 2–4x faster with 10–20x less memory — and now runs inside nearly every major LLM.
- The insight: attention's real bottleneck wasn't the quadratic math but GPU memory traffic; FlashAttention is an exact algorithm redesigned around the memory hierarchy.
- It enabled longer context windows at the same cost, and was quickly integrated into PyTorch and Hugging Face.
- The arc: FlashAttention-2 (July 2023, 2x faster), Mamba with Albert Gu (rejected by ICLR 2024, still hugely influential), FlashAttention-3 for Hopper, plus per-generation chip tuning. Dao is now a Princeton professor and Together AI chief scientist who still writes GPU kernels himself.
More from Infra
- Anthropic reportedly buying ~5GW of compute from Broadcom, may become its largest customer by 2027 — zephyr_z9 · 2026-10-01
- Supertonic: open-source 99M-param local TTS beats ElevenLabs in tests — JafarNajafov · 2026-10-01
- Zoho opens its self-built cloud as Catalyst serverless platform, free for students — RoboBalaji · 2026-10-01
- Running vLLM on mixed NVIDIA 50+40 GPUs: a DIY resource config repo — Fz1zz · 2026-10-01
- Diffusers tensor parallel loading gets 2.4x faster, cuts CPU memory by 89% — RisingSayak · 2026-10-01
- Building an RTX 6000 cluster: ducting exhaust to the wall and flipping 6000 Pros to spare the motherboard — TheZachMueller · 2026-10-01