Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration
illinois · hf · 2026-09-30
University of Illinois researchers study the overlooked softmax in low-precision Transformer training. Their K-interval attention approximates the exponential with K+1 grid values, ablating per-row grid calibration, interpolation vs hard rounding, and straight-through surrogate placement, with derived backward rules including calibration derivatives.
On matched pretraining runs (124M params, 2.5B tokens):
- At K=4 with hard rounding, min-max calibration plus a pre-normalization surrogate shows a large loss gap; changing either choice substantially reduces it.
- Fixed-window calibration with a post-normalization surrogate yields +0.019 nats validation loss gap at K=4; with a pre-normalization surrogate only +0.004 nats at K=16.
- Detaching row extrema leaves forward computation unchanged but causes a delayed validation-loss increase.
More from Infra
- Where should neolabs get GPUs? Insider votes put SF Compute ahead of CoreWeave — FinanceYF5 · 2026-09-30
- Google donates Agent Substrate to CNCF as runtime layer for large-scale agents — rakyll · 2026-09-30
- WSL Containers is now generally available: run Linux containers on Windows via wslc CLI — solyarisoftware · 2026-09-30
- South Korea expects record-high tax revenue amid AI chip boom — Polymarket · 2026-09-30
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30
- Ubuntu Snapdragon Edition ISO now available for download ahead of official 2027 release — carrycooldude · 2026-09-30