Benchmarking Qwen2.5-32B Quantization: Custom AD-IQ3_S Beats Community by 33%
Top-Eye-8104 · reddit · 2026-08-15
The team quantized Qwen2.5-32B and benchmarked it against 20 community GGUF files (from Unsloth, LMStudio, etc.) using a consistent setup on 4x RTX 5090s.
Key Findings:
- Performance: Between 12GB and 21GB, their quantization curves show lower drift (better accuracy) than any community version. At 13.8GB, their AD-IQ3S shows 33% less drift than Unsloth's Q3KM.
- Extremes: Performance degrades rapidly for all quants below 10GB; differences above 25GB are negligible.
- Recommendation: For 16GB hardware, AD-IQ3S (13.8GB) is the top pick with 92.4% top-1 accuracy.
- Release: Weights, Imatrix, and layouts are available on Hugging Face.
More from Infra
- Nvidia close to deal to guarantee about $100B in credit for OpenAI — pstAsiatech · 2026-08-15
- AlphaSense Study: Context is the Bottleneck, GPT-5.6 Beats Kimi on Cost Efficiency — rohanpaul_ai · 2026-08-15
- Study: Anthropic models may be cheaper than some open-source Chinese models — rohanpaul_ai · 2026-08-15
- Polygres turns Postgres into extended context for AI agents — Scobleizer · 2026-08-15
- vLLM Introduces Adaptive Verification for Speculative Decoding with DSpark — vllm_project · 2026-08-15
- Open source closes the gap with closed labs: Quality gap now just months — togethercompute · 2026-08-15