Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration

illinois · hf · 2026-09-30

University of Illinois researchers study the overlooked softmax in low-precision Transformer training. Their K-interval attention approximates the exponential with K+1 grid values, ablating per-row grid calibration, interpolation vs hard rounding, and straight-through surrogate placement, with derived backward rules including calibration derivatives.

On matched pretraining runs (124M params, 2.5B tokens):

Original post →

More from Infra

Infra channel →