Does DIY quantizing Qwen 3.8 still make sense when Unsloth UD 3.0 exists?
Professional-Tap177 · reddit · 2026-08-31
- The author recreated the community's smaller Qwen3.8-27B IQ4XS GGUF, swapping in the Unsloth imatrix, and published it on Hugging Face.
- Core question: with Unsloth Dynamic (UD) 3.0 quantization being notably better, is hand-rolling quantization with llama-quantize still worthwhile?
- Not yet benchmarked, but the author suspects it won't beat the UD 3.0 Q3 quants.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01
- Data Center Worker: Fastest Blue-Collar Path to Six Figures Right Now — AICopyLab · 2026-09-01