New Quantization Method Boosts Quality on Qwen3.5

1ncehost · reddit · 2026-07-12

The author has released two new sets of GGUF quantization results based on Qwen3.5 0.8B / 2B, claiming that their new method, Voodoo Quant, outperforms Unsloth Dynamic 2.0 KLD in mixed-precision optimization.

The post emphasizes that similar to Unsloth Dynamic, this approach allocates higher precision to more critical parts of the model. However, Voodoo Quant optimizes across the entire tensor rather than block by block. The author also compares KLD performance between Torch and llama.cpp graph structures, suggesting that while Unsloth performs well in llama.cpp, it degrades noticeably on Torch graphs, indicating a potential overfitting to llama.cpp. In contrast, Voodoo offers a more balanced performance across both. The author believes it excels at more aggressive quantization levels, even suggesting that 2 bit might be the sweet spot.

Related event: New Voodoo Quant Claims to Surpass SOTA(3 posts)→

Original post →

More from Infra

Infra channel →