Undergrad's improved GPTQ beats official AWQ on 4-bit Qwen2.5 perplexity
Status-Adeptness8123 · reddit · 2026-10-05
An undergrad built an improved GPTQ quantizer in two weeks on a MacBook using Claude as a coding assistant, with results reproduced independently on an A10G. Three additions to vanilla GPTQ: per-group grid fitting instead of min-max, a second pass re-checking every rounded weight, and grid refitting against layer input statistics. Integer zero points let models pack into standard AWQ format and run on vLLM's awqmarlin kernel.
WikiText-2 perplexity / HumanEval pass@1 (A10G, vLLM 0.29):
- Qwen2.5-1.5B: fp16 9.37/37.2% → custom 9.66/33.5% vs official AWQ 10.16/34.1%
- Qwen2.5-7B: fp16 7.15/70.1% → custom 7.29/67.1% vs official AWQ 7.58/64.6%
Honest caveats: single runs, HumanEval deltas within noise (3.6 pts), calibration on WikiText-2 favors the perplexity test, the lead shrinks at larger scale, and at 3 bits code/math ability drops by more than half. Code, models (incl. MLX) and failed experiments are open-sourced.
More from Infra
- XFreeze: high-bandwidth memory holders will win the superintelligence race in 1-3 years — XFreeze · 2026-10-06
- 691K H100-hours: community estimates compute behind from-scratch checkpoint trained on just 320 H100s — teortaxesTex · 2026-10-06
- Dev reports Clef-Flash runs fast locally even on memory-bandwidth-limited Jetson Orin — gregmushen · 2026-10-06
- Running LLMs Fully Client-Side: llama.cpp Compiled to WASM with WebGPU Proof of Concept — Numerous-Fan8138 · 2026-10-06
- Llama.wasm: llama.cpp Compiled to WASM with WebGPU Runs LLMs Fully in the Browser — Numerous-Fan8138 · 2026-10-06
- NVIDIA Dynamo lets coding agents point at self-hosted endpoints with native tracing — TheZachMueller · 2026-10-06