Benchmarked 21 Qwen3.8-27B quants on 16GB VRAM: bartowski IQ4_XS wins
Storterald · reddit · 2026-09-05
A Reddit user benchmarked 21 quantized variants of Qwen3.8-27B on an RTX 5080 (16GB) using their own C code, ranked by Mean KLD (output distribution divergence from the original model).
Key findings:
- Best overall: bartowski/Qwen3.8-27B-IQ4XS (KLD 0.056, 14.5GiB, fits in 16GB)
- Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4XS (KLD 0.083)
- For smaller footprints: jpetrina's IQ4XS-pure or unsloth's UD-Q3KXL (KLD 0.14, just over 12GiB)
- unsloth UD-Q4KXL (KLD 0.028, 16.4GiB) is the quality ceiling but doesn't fit
- 2-bit quants (sdkyuan QAT Q20, ISTA-DASLab GSQ-RCO IQ2 series) degrade significantly
Overall: 4-bit is nearly lossless, 3-bit quite usable, 2-bit noticeably worse.
More from Infra
- Notion's AI Meeting Notes run on Baseten's voice inference stack — baseten · 2026-09-05
- Two-person team hits SOTA in GPU failure prediction with a $2 PINN algorithm — sharpeye_wnl · 2026-09-05
- Huawei to produce under 4% of Nvidia's AI compute in 2026, Epoch AI estimates — Jsevillamol · 2026-09-05
- Hugging Face downloads crawl at 700kb/s on 1Gbit lines, users float p2p model distribution — perelmanych · 2026-09-05
- After ChatGPT, Claude & Grok All Went Dark, One User's 3-Machine Local AI Lab Kept Running — cocktailpeanut · 2026-09-05
- DRAM density has flattened: servers stuck at 8TB for 5 years, CXL is the way out — lauriewired · 2026-09-05