G²PTQ: generalized gradient compensation improves LLM post-training quantization

Ruikang Liu · hf · 2026-09-29

G²PTQ is a unified post-training quantization (PTQ) framework addressing two complementary flaws in GPTQ-style methods: local layer-wise objectives lack global supervision, while global-objective methods fix Hessian estimates upfront, ignore first-order gradients, and go stale as quantization proceeds.

Method

Results: better alignment with full-precision models across model families and bit-widths, outperforming SOTA baselines; code released on GitHub.

Original post →

More from Research

Research channel →