Unsloth Releases Qwen3.8 Local Deployment Guide with 1-bit Quantization Down to 397GB
MaziyarPanahi · x · 2026-08-13
Unsloth has released a local deployment guide for the Qwen3.8 model family. The series includes 27B, 2.4T-A95B (2.4T total/95B active params), and Max versions, featuring vision and thinking capabilities with up to a 1M token context window.
To handle the massive size, Unsloth extended llama.cpp's quantization with new 1-bit data types (e.g., UD-IQ1XXXS). By reducing codebook entries to 1.1875 bits per weight, the 2.4T model is compressed to 397GB (91% smaller) for local execution, retaining strong accuracy without requiring Quantization-Aware Training (QAT). The 27B model runs on 16GB+ memory setups.
Related event: Unsloth Shrinks Qwen3.8 by 91% for Local Deployment(3 posts)→
More from Infra
- Open-Source mlx-dspark Boosts LLM Inference on Mac by 3.3x — A-Rahim · 2026-08-13
- Fluidstack at NYSE: Discussing Capital Behind US AI Infrastructure — MxMnr · 2026-08-13
- Open-Source CUDA Alternative for Portable AMD GPU Code — tom_doerr · 2026-08-13
- Slow TPS on AMD 7900 XT Running Gemma: 131K Context Bottleneck — opoot_ · 2026-08-13
- Building Local Open-Weight Agents on a 4GB VRAM GPU — ComplexHuman26 · 2026-08-13
- Polymarket Indicates Only a 15% Chance of an AI Bubble Burst by Year-End — Polymarket · 2026-08-13