Unsloth Releases Qwen3.8 Local Deployment Guide with 1-bit Quantization Down to 397GB

MaziyarPanahi · x · 2026-08-13

Unsloth has released a local deployment guide for the Qwen3.8 model family. The series includes 27B, 2.4T-A95B (2.4T total/95B active params), and Max versions, featuring vision and thinking capabilities with up to a 1M token context window.

To handle the massive size, Unsloth extended llama.cpp's quantization with new 1-bit data types (e.g., UD-IQ1XXXS). By reducing codebook entries to 1.1875 bits per weight, the 2.4T model is compressed to 397GB (91% smaller) for local execution, retaining strong accuracy without requiring Quantization-Aware Training (QAT). The 27B model runs on 16GB+ memory setups.

Related event: Unsloth Shrinks Qwen3.8 by 91% for Local Deployment(3 posts)→

Original post →

More from Infra

Infra channel →