BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training
HongyiWang10 · x · 2026-10-04
A hidden gradient bug in BF16 attention
A Rutgers team pretrained a 450M-parameter transformer on 50B tokens with FlashAttention-3 in BF16. Training was healthy for 25B tokens, then the gradient norm grew 1000× and the final loss ended 0.2 nats above FP32 attention — without a single NaN. Healthy-looking training does not mean correct gradients.
- Cause 1: a fused multiply-add rounding issue in the forward softmax (previously treated as an extreme-input NaN case, never fixed in FlashAttention-3).
- Cause 2: BF16 rounding breaks the conservation law that softmax score gradients sum to zero along each row, leaking the mean key into the query gradient — a leak that grows exactly as late training makes keys large and attention sharp.
Fix: GProj
The authors propose GProj (gauge projection), restoring the zero-sum law after rounding. Query/key gradient accuracy and final loss match FP32 attention with only 4.7% extra step time. Recomputing just two layers' attention backward in FP32 removes nearly all excess gradient.
> Training looks fine? Check the backward pass anyway.
More from Infra
- Datacenter buildout would have been less aggressive in a world where everyone gets AGI at once — willcb · 2026-10-04
- UBS Sees Global Rack Capacity Hitting 104.5 GW by 2030, Half Going to Nvidia — AccBalanced · 2026-10-04
- Musk: xAI to hit 10GW of compute by end of next year, sees AI inference moving to space — beffjezos · 2026-10-04
- Getting PyTorch CUDA training running on BC-250 boards, captured as an image — redfoxkiller · 2026-10-04
- Huawei 950 super-node claims seamless scaling from 550B/1.6T up to 10T-class models — teortaxesTex · 2026-10-04
- The curse of 64GB RAM: Strata pushes local Qwen3.8-Flash-Next to 60 t/s but hogs system memory — Cautious_Chicken_604 · 2026-10-04