KohakuFA: Blackwell flash attention kernel fixes 10x-1000x gradient bugs in FA4/cuDNN
bdsqlsz · x · 2026-10-04
KBlueleaf's team found that mainstream attention kernels — cuDNN, FlashAttention-4, and flexattention — all suffer severe backward-pass precision issues: once an attention layer learns near-one-hot attention and logits exceed 1e5, gradients are wrong by 10x to 1000x (or inf) while forward and loss look normal. Many "unstable fp16 training" failures were the kernel's fault, not fp16's. They released KohakuFA, a flash attention implementation for NVIDIA Blackwell (sm100+) written in Gluon (Triton's tile-level dialect) using tcgen05 MMAs, TMA, and warp specialization, supporting dense, token-causal, and block-causal attention. Its fp16 gradients stay at the fp16 rounding floor, close to fp64 and FlashAttention-2. Details in docs/precision.md.
More from Infra
- Anthropic's Sholto Douglas: AI capex could hit $4T by 2028 and double GDP by early 2030s — ziv_ravid · 2026-10-04
- Pretraining Still Matters: RL Is Just More Compute-Intensive, Not More Important — akbirthko · 2026-10-04
- Bain: hyperscalers now planning 5GW data center campuses costing $150-200 billion each — Beth_Kindig · 2026-10-04
- 150-300 H100s available for short-term burst rental with IB and fast local storage — AccBalanced · 2026-10-04
- Compute shortage may force second-tier neoclouds to diversify beyond NVIDIA chips — AccBalanced · 2026-10-04
- Running a 176B MoE on a 16GB RTX 3080 laptop: TensorSharp beats Strata in end-to-end test — fuzhongkai · 2026-10-04