HyQuant: hybrid-precision quantization keeps critical attention tokens in high precision

SJTU · hf · 2026-09-11

HyQuant from SJTU improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.

Original post →

More from Research

Research channel →