Unsloth Releases Qwen2.5-72B Quants: 1-bit Version Runs on 8GB RAM
cephaloform · x · 2026-08-20
Unsloth AI has released new GGUF quantizations for Qwen2.5-72B. The Dynamic V3 version achieves >10% higher accuracy on benchmarks like Div-300 and KLD. Notably, they also released a 1-bit extreme quantization that runs on just 8GB RAM while retaining approximately 77% of BF16 accuracy.
More from Infra
- Monad Agent Hub launches with no-code platforms for instant agent creation — bgmshana · 2026-08-20
- AWS Leverages AI Infrastructure Demand to Extend Cloud Dominance — DavidLinthicum · 2026-08-20
- Reverse-Engineering RK3588 NPU: Open Compiler Runs GPT-2 at 36 tok/s — one_does_not_just · 2026-08-20
- llama.cpp dflash2: Qwen 3.8 27B Inference Speed Up to 3x — Top-Eye-8104 · 2026-08-20
- NVIDIA announces FLARE Day event for September — AllThingsApx · 2026-08-20
- Dev billing guide: Pay-per-token may undercut subscriptions — heypearlai · 2026-08-20