Run Qwen3.8 Locally: Unsloth Shrinks Model Size by 91% via 1-bit Quantization
danielhanchen · x · 2026-08-13
Unsloth AI announced that they have successfully shrunk the massive Qwen3.8-2.4T-A95B model from 4.9TB to 397GB (a 91% reduction) using their new Dynamic 1-bit quantization, enabling local execution.
- Quantization Breakthrough: By selectively quantizing layers and reducing codebook entries (down to 1.1875 bits per weight), the method retains accuracy while drastically cutting size. It works well for post-training quantization (PTQ) without needing QAT or QAD.
- Hardware Requirements: The compressed model can run on setups with 410GB+ RAM/VRAM via Unsloth Desktop or llama.cpp.
- Model Capabilities: Qwen3.8 features vision and thinking capabilities with a 256K context window (up to 1M tokens), rivaling GPT-5.6 Sol.
Related event: Unsloth Enables Local Qwen3.8 Deployment via 1-bit Quantization(2 posts)→
More from Infra
- SkyPilot Unifies Multiple Slurm Clusters, Solving GPU Management Bottlenecks — skypilot_org · 2026-08-13
- 2-bit Quantized Nemotron 3.5 Runs Autonomous Tool Calls Continuously on Just 22GB VRAM — danielhanchen · 2026-08-13
- Report: SpaceXAI Builds Custom GB300 Inference Stack for 2× Performance Gains — XFreeze · 2026-08-13
- Tracking Amazon Bedrock Costs with Athena and CUDOS Dashboards — AWS ML Blog · 2026-08-13
- Running Minimax H3 Locally: A 6GB RTX 3050 VRAM Test — JadedScorpion · 2026-08-13
- AI Boom Spreads: Investors Target Chip Fab and Data Center Suppliers — Polymarket · 2026-08-13