Rowmax-H15: approximate softmax in attention for 25.8% faster B200 inference with minimal quality loss

illinois · hf · 2026-09-30

Researchers characterize how far softmax can be approximated at inference across 10 frozen decoder-only models (0.5B–72B).

Findings:

Proposed:

Measured on B200: FP8 attention forward is 12.4% faster at causal 8K and 25.8% at non-causal 8K; board energy drops 8.4% at causal 16K; BF16 perplexity rises only 0.091–0.492% across five models.

Original post →

More from Infra

Infra channel →