FlashAttention shows why FLOPs alone are a bad lens on GPU performance
techNmak · x · 2026-09-10
Using FlashAttention as a case study, the author explains why FLOPs are a poor proxy for GPU performance. Standard dense attention (QK^T, softmax, times V) is already highly parallel — the real pain is what naive implementations do around the math: materializing huge N×N intermediates and round-tripping them through high-bandwidth memory.
FlashAttention's core contribution is optimizing exactly those memory accesses: tiling plus online softmax keeps intermediates on-chip in SRAM instead of hitting HBM, dramatically improving real-world throughput. A great illustration that attention is memory-bound, not compute-bound.
Related event: Understanding FlashAttention: Why Memory Traffic Beats FLOPs(2 posts)→
More from Infra
- Positron AI Raises $875M Series C at $5B Valuation, Deploying 50+ Atlas Racks at Oracle Cloud — Scobleizer · 2026-09-11
- Positron AI raises $230M Series B at over $1B valuation with Arm backing — seanmcdonaldxyz · 2026-09-11
- Cerebras Fast Inference Flips Agent Workflows: Fewer Parallel Agents, Same Output — MatthewBerman · 2026-09-11
- Baseten acquires Blaxel to build integrated cloud infrastructure for AI agents — baseten · 2026-09-11
- Skild AI's S1 learns robot tasks from one video, hits $100M revenue run rate — NVIDIA Blog · 2026-09-11
- Hyperscalers could factor RSA-1024 for about $30M per number, analysis claims — rickasaurus · 2026-09-11