Undergrad's improved GPTQ beats official AWQ on 4-bit Qwen2.5 perplexity

Status-Adeptness8123 · reddit · 2026-10-05

An undergrad built an improved GPTQ quantizer in two weeks on a MacBook using Claude as a coding assistant, with results reproduced independently on an A10G. Three additions to vanilla GPTQ: per-group grid fitting instead of min-max, a second pass re-checking every rounded weight, and grid refitting against layer input statistics. Integer zero points let models pack into standard AWQ format and run on vLLM's awqmarlin kernel.

WikiText-2 perplexity / HumanEval pass@1 (A10G, vLLM 0.29):

Honest caveats: single runs, HumanEval deltas within noise (3.6 pts), calibration on WikiText-2 favors the perplexity test, the lead shrinks at larger scale, and at 3 bits code/math ability drops by more than half. Code, models (incl. MLX) and failed experiments are open-sourced.

Original post →

More from Infra

Infra channel →