Flash Attention Should Be Default Over SDPA
LysandreJik · x · 2026-07-15
This post emphasizes that **there's no need to use SDPA attention when running LLMs**, equating it to "handicapping your own hardware." The author argues that while deploying flash attention used to be a hassle, it is no longer a major barrier. Key takeaways: - **flash attention** has a significant impact on throughput; - Continuing to use **SDPA** needlessly sacrifices performance; - This optimization is mature enough today that it shouldn't be treated as an "optional feature." The original post also mentions that the author didn't even want to expand on the performance gains brought by **CB** (further performance optimization topics in context).
More from Infra
- Larry Fink says China is ahead in the AI energy race, citing 100 GW nuclear buildout — rohanpaul_ai · 2026-07-21
- Local AI may pay back in 6–7 years and cut long-term costs by 30–40% — DavidLinthicum · 2026-07-21
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21
- More open models and llama.cpp updates are coming, says Merve Noyan — mervenoyann · 2026-07-21
- Why adding a second LLM provider breaks more than the API surface — Ok_Extension6373 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21