G²PTQ: generalized gradient compensation improves LLM post-training quantization
Ruikang Liu · hf · 2026-09-29
G²PTQ is a unified post-training quantization (PTQ) framework addressing two complementary flaws in GPTQ-style methods: local layer-wise objectives lack global supervision, while global-objective methods fix Hessian estimates upfront, ignore first-order gradients, and go stale as quantization proceeds.
Method
- Integrates first- and second-order information under a globally supervised, block-wise objective
- Refreshes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness
- A trust-region scaling mechanism dynamically bounds the gradient step to prevent exploding updates
- Efficient implementations for block-wise Hessian approximation and exact gradient compensation
Results: better alignment with full-precision models across model families and bit-widths, outperforming SOTA baselines; code released on GitHub.
More from Research
- Judea Pearl recommends Bareinboim's talk on causal AI in the age of LLMs — yudapearl · 2026-09-29
- The overlooked async RL curation pitfall: task runtime differences skew sampling — auto_grad_ · 2026-09-29
- RLDM 2027 heads to Paris, July 6-9, with DeepMind's David Abel as program chair — dabelcs · 2026-09-29
- SNaP one-step posterior sampling runs 30x faster than iterative samplers — prof_kamilov · 2026-09-29
- No matched data? Climate study's two-stage pattern cuts humid-heat bias by 44-57% — bravo_abad · 2026-09-29
- MIT Team's AI Agents Uncover Design Principles That Keep Atom-Thin Graphene From Catastrophic Failure — ProfBuehlerMIT · 2026-09-29