Rowmax-H15: approximate softmax in attention for 25.8% faster B200 inference with minimal quality loss
illinois · hf · 2026-09-30
Researchers characterize how far softmax can be approximated at inference across 10 frozen decoder-only models (0.5B–72B).
Findings:
- On B200, tensor-core throughput outpaces special-function exponential throughput by 2+ orders of magnitude, making exp evaluation the bottleneck—yet pretrained models don't need it exact everywhere.
- Probability-assigned positions and within-row resolution can be cut substantially, but uniform weighting of the same positions is damaging; resolution budget near the row max is consistently favored; identical scalar distortions produce opposite-sign responses across models.
Proposed:
- Rowmax-PoT: a coarse logarithmic weight representation anchored at each row max;
- Rowmax-H15: hardware specialization in FlashAttention-4.
Measured on B200: FP8 attention forward is 12.4% faster at causal 8K and 25.8% at non-causal 8K; board energy drops 8.4% at causal 16K; BF16 perplexity rises only 0.091–0.492% across five models.
More from Infra
- DeepSeek open-sources DeepGEMM Ascend port, hitting 99.8% of hardware limit on GEMM — zheanxu · 2026-09-30
- AI intelligence-cost Pareto frontier shifted fast: GPT-5 mini at 17 ($0.05) to Claude Opus 5.5 at 58 ($5.98) — ArtificialAnlys · 2026-09-30
- Google's Project Suncatcher to fly TPUs in space for the first time on Oct 1 — allisondman · 2026-09-30
- Photon 2.6 ships FP8 + speculative decoding, runs Qwen3.5 27B at 400+ tok/s on B200 — Bedrovelsen · 2026-09-30
- SGLang turns Qwen3.8-27B into a decision model that beats Pokémon FireRed at sub-100ms — zhaoran_wang · 2026-09-30
- Agentic AI turns CPUs into the overlooked bottleneck as CPU:GPU ratios shift upward — AccBalanced · 2026-09-30