Adobe's softmax reparameterization cuts output-head W4 quantization KL error by 73%
adobe · hf · 2026-09-28
Adobe researchers propose softmax reparameterization, a post-training method for output-head quantization in small LLMs.
- Idea: Subtract a scalar multiple of the vocabulary-row mean from every output row to pick a functionally equivalent head before quantization; the coefficient is chosen by a one-dimensional validation-KL search per quantizer (RTN, activation-weighted MSE, full-Hessian GPTQ).
- Properties: Preserves the full-precision softmax distribution, leaves the decoder unchanged, and a rank-one correction handles nonlinear logit paths like soft-capping.
- Results: On Phi-4-mini, AW-MSE KL drops from 0.936 to 0.256 at W4; gains are complementary to per-channel scaling and affine quantization, and frozen WikiText-selected coefficients transfer to C4 and OpenWebText, outperforming mean-centering in 18 of 24 comparisons.
- Inference: Adds no operations for shift-compatible heads and cuts batch-one generation latency by 10.8% with a BF16 decoder; benefits broaden at W2.
More from Research
- Goodfire grants geometric_intel lab funding for AI interpretability research — ninamiolane · 2026-09-29
- Bespoke Labs Launches AutoResearchExam Benchmark, Again, With a Demo Video — gregd_nlp · 2026-09-28
- Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner — pcastr · 2026-09-28
- IROS 2026 has 1,933 papers — researcher curates 130-paper reading list on VLA and robot learning — GlenBerseth · 2026-09-28
- Ex-game-AI developer: general agents are taking over bespoke game AI systems — weballergy · 2026-09-28
- Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf — weballergy · 2026-09-28