New Quantization Method Boosts Quality on Qwen3.5
1ncehost · reddit · 2026-07-12
The author has released two new sets of GGUF quantization results based on Qwen3.5 0.8B / 2B, claiming that their new method, Voodoo Quant, outperforms Unsloth Dynamic 2.0 KLD in mixed-precision optimization.
The post emphasizes that similar to Unsloth Dynamic, this approach allocates higher precision to more critical parts of the model. However, Voodoo Quant optimizes across the entire tensor rather than block by block. The author also compares KLD performance between Torch and llama.cpp graph structures, suggesting that while Unsloth performs well in llama.cpp, it degrades noticeably on Torch graphs, indicating a potential overfitting to llama.cpp. In contrast, Voodoo offers a more balanced performance across both. The author believes it excels at more aggressive quantization levels, even suggesting that 2 bit might be the sweet spot.
Related event: New Voodoo Quant Claims to Surpass SOTA(3 posts)→
More from Infra
- NVIDIA publishes Vera CPU architecture details before AMD’s AI event — ryanshrout · 2026-07-22
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- HilbertRaum open-sources a fully local AI chat and document analysis app for private use — Vladowski · 2026-07-22
- Hybrid and local inference are emerging as a response to AI energy and token costs — dmitry140 · 2026-07-22