HyQuant: hybrid-precision quantization keeps critical attention tokens in high precision
SJTU · hf · 2026-09-11
HyQuant from SJTU improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.
More from Research
- CiNet's 11th international conference on biological to general AI opens registration — kaixhin · 2026-09-11
- Looped Transformers Hit a Natural Ceiling: Most Test-Time Scaling Is Environment Interaction — generativist · 2026-09-11
- New attack reconstructs local LLM outputs from CPU cache traces, up to 95% success — rohanpaul_ai · 2026-09-11
- Ryan Greenblatt questions whether plain-text CoTs will really become an alchemical era of observability — RyanGreenblatt · 2026-09-11
- Classic free book 'From Python to NumPy' now readable online with an AI tutor — burkov · 2026-09-11
- Author's caveat: stroke model's Dice 0.46 is an upper bound from outcome-fitted parameters — maier_ak · 2026-09-11