PrismQuant: null-space rotations make INT4 near-lossless, only 0.22pp below FP16 on Llama-70B
NanyangTechnologicalUniversity · hf · 2026-09-30
PrismQuant aligns activation eigenspaces with the constant group subspace of asymmetric grouped INT4 via a closed-form Ky Fan trace-maximizing rotation built from compact Householder transforms. Under W4A4KV4 it sets SOTA on Llama-3.2-3B; Llama-3.1-70B reaches 3.85 perplexity and 72.46% zero-shot accuracy (only 0.22pp below FP). Deployed on Llama-3.1-8B it yields 1.51x prefill and 1.22x decode speedups with 56.34% lower peak decode memory. Code released.
More from Infra
- 8% of Asia-to-US air freight is now data center parts — 30 full freighters a day — yacineMTB · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- Hugging Face ships tokenizers v1, often tens of times faster than v0.23 — ariG23498 · 2026-09-30
- Altman says OpenAI's custom chip program comes online in H1 2027, betting on inference advantage — johncoogan · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- AAOI burnt capex on US vertical integration, CW laser yield still poor — jwt0625 · 2026-09-30