BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training

HongyiWang10 · x · 2026-10-04

A hidden gradient bug in BF16 attention

A Rutgers team pretrained a 450M-parameter transformer on 50B tokens with FlashAttention-3 in BF16. Training was healthy for 25B tokens, then the gradient norm grew 1000× and the final loss ended 0.2 nats above FP32 attention — without a single NaN. Healthy-looking training does not mean correct gradients.

Fix: GProj

The authors propose GProj (gauge projection), restoring the zero-sum law after rounding. Query/key gradient accuracy and final loss match FP32 attention with only 4.7% extra step time. Recomputing just two layers' attention backward in FP32 removes nearly all excess gradient.

> Training looks fine? Check the backward pass anyway.

Original post →

More from Infra

Infra channel →