Unsloth Releases Kimi K3 Quantized Weights, 1-bit Compresses to 594GB
The Unsloth team has released a quantized version of the Kimi K3 model, enabling this ultra-large model to run on local devices. Notably, the 1-bit quantized version compresses the original 1.56TB model down to 594GB while retaining 78.9% accuracy, presenting a highly feasible solution for local deployment of massive AI models.
已确认
- Kimi K3 is an open-weight model featuring 2.8T parameters and 104B active parameters, equipped with native vision capabilities and a 1 million token context window.
- Unsloth has uploaded various quantization schemes for the model to Hugging Face, including a 1.5TB MXFP4 version and GGUF packages, along with multimodal projection files.
- Through 1-bit quantization, the model size is reduced from 1.56TB to 594GB (a 62% reduction) while maintaining approximately 78.9% accuracy.
- The quantized Kimi K3 has been successfully run locally on high-memory Mac Studio configurations.
为什么重要
- 降低部署门槛: Compressing a 1.5TB+ ultra-large model to under 600GB means everyday developers and research institutions can run top-tier LLMs locally using consumer or prosumer hardware (like the Mac Studio), breaking the absolute reliance on expensive cloud compute.
- 验证量化技术: Retaining nearly 80% accuracy despite a massive 62% size reduction via 1-bit quantization proves the practical potential of extreme quantization techniques for ultra-large models.
2026-07-28 ~ 2026-07-30 · 6 related posts
Primary sources
- Unsloth Releases Kimi K3 GGUFs: MXFP4 Version Reaches 1.5TB — _TheWolfOfWalmart_ ·
- Unsloth Releases Kimi K3 Quantized: 1-bit Compresses to 594GB Retaining 78.9% Accuracy — BankApprehensive7612 ·
- Unsloth Releases Kimi K3 Quantized Version — _akhaliq · 2026-07-28
- [source] Unsloth Releases Kimi K3 GGUFs: MXFP4 Version Reaches 1.5TB — _TheWolfOfWalmart_ · 2026-07-29
- Moonshot’s Kimi K3 now runs locally at 594 GB with 78.9% top-1 accuracy — danielhanchen · 2026-07-29
- Quantized Kimi K3 lands on Hugging Face in GGUF formats — victormustar · 2026-07-29
- 1-bit Quantization Magic: Kimi K3 Successfully Runs Locally on Mac Studio — danielhanchen · 2026-07-30
- [source] Unsloth Releases Kimi K3 Quantized: 1-bit Compresses to 594GB Retaining 78.9% Accuracy — BankApprehensive7612 · 2026-07-30