Unsloth Dynamic V3 GGUFs: Q3 Outperforms Larger Q4 Models
danielhanchen · x · 2026-08-21
Tests of Unsloth's new Dynamic V3 GGUFs on a 4090 show their Q3 quantization outperforms a Q4 model that is 4.5GB larger.
Benchmark results (Wikitext-2, 262k ctx):
- Unsloth UD-Q3: 12.24GB, PPL 6.3993, Median 110.7 t/s, Retrieved needle at 250k tokens.
- Comparison: The standard Q4 was larger (16.7GB), slower (62.6 t/s), and only retrieved the needle at 4k tokens.
The performance comes from layer-specific types derived from error analysis. Unsloth also released 1-bit quants that retain 77% accuracy and run on 8GB RAM.
More from Infra
- Investor poll ranks AI supply chain startups under $10B: MatX leads, Lightmatter and Modal close — FinanceYF5 · 2026-08-21
- Americans Prefer Coal Plants Over Data Centers: Study — Polymarket · 2026-08-21
- OpenBMB releases Ultra-FineWeb-L1, a 1T+ token high-quality web dataset — zibuyu9 · 2026-08-21
- AI Data Centers Bypass Grid Delays with On-Site Gas and Solar — shensi · 2026-08-21
- VC Trend: Invest in Own GPU Clusters to Offer Compute at Cost to Portfolio Companies — beffjezos · 2026-08-21
- Heterogeneous Systems Outperform Frontier Models in New Benchmarks — ShahabBakht · 2026-08-21