KohakuFA: Blackwell flash attention kernel fixes 10x-1000x gradient bugs in FA4/cuDNN

bdsqlsz · x · 2026-10-04

KBlueleaf's team found that mainstream attention kernels — cuDNN, FlashAttention-4, and flexattention — all suffer severe backward-pass precision issues: once an attention layer learns near-one-hot attention and logits exceed 1e5, gradients are wrong by 10x to 1000x (or inf) while forward and loss look normal. Many "unstable fp16 training" failures were the kernel's fault, not fp16's. They released KohakuFA, a flash attention implementation for NVIDIA Blackwell (sm100+) written in Gluon (Triton's tile-level dialect) using tcgen05 MMAs, TMA, and warp specialization, supporting dense, token-causal, and block-causal attention. Its fp16 gradients stay at the fp16 rounding floor, close to fp64 and FlashAttention-2. Details in docs/precision.md.

Original post →

More from Infra

Infra channel →