SOTA GGUF quants for Qwen3.8-Flash-Next: near-baseline quality at 1/5 the size
BullfrogScary8947 · reddit · 2026-09-16
ISTA-DASLab released GSQ-RCO GGUF quantizations of Qwen3.8-Flash-Next, cutting the 80-95GB model to 68-76GB with near-baseline quality. Key points:
- IQ3XXS is the strongest point: matches baseline exactly on AIME25 (100.00), within 0.51 on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of BF16 size.
- Q20 targets speed: avoiding lookup-table-heavy formats, it delivers 3.4x prompt throughput and 1.9x lower end-to-end latency vs IQ2XS, with 6.2x better prompt throughput in coding; decode rate stays flat across workloads. Trade-off: 89.07 vs 89.16 task average, 3.5 points below IQ3XXS.
Available on Hugging Face for local deployment, pick per throughput/quality needs.
More from Infra
- Microsoft bets on local AI: Windows agent stack spans $800 Copilot+ PCs to 1T-param DGX Station — ryanshrout · 2026-09-16
- PlanetScale Traffic Control lets you budget DB resources per app name — DanielLockyer · 2026-09-16
- GPUs as VC value-add: European AI startups' top constraint is compute access — nellimorgulchik · 2026-09-16
- Community squeezes a 124B model onto a 128GB DGX Spark with quantization and kernel fixes — alifcoder · 2026-09-16
- One chart explains how CPU, GPU and TPU differ — mdancho84 · 2026-09-16
- Even 4-year-old GPUs are repricing up: 20% renewal premium, contracts up 125% — tengyanAI · 2026-09-16