1-bit Quantized Qwen 27B Runs on 12GB VRAM at 92 tokens/s
zyxciss · reddit · 2026-08-16
Reddit user zyxciss tried a 1-bit quantized version of Qwen3.8-27B (AtomicChat/Qwen3.8-27B-GGUF on HuggingFace) and managed to run it on an RTX 3060 12GB. With MTP enabled, it achieves about 92 tokens/s.
Although the model quality is severely degraded (currently unusable), the speed on a 3060 is surprising. The user had hoped for a 35B MoE version, but that seems unlikely.
More from Infra
- llama.cpp integrates Dots3 Note model, scoring 78.4 on SWE-bench Verified — victormustar · 2026-08-16
- Weaviate adds test-time compute scaling to Search Mode, boosting retrieval performance significantly — dl_weekly · 2026-08-16
- New Book: Algorithms for Modern Hardware Open-Sourced on GitHub — thehiphopswami · 2026-08-16
- Comfy Kitchen Attention Speeds Up MiniMax H3 on AMD GPUs — God_Hand_9764 · 2026-08-16
- Quality difference between Q8_0 and UD-Q6_K_XL quantization — AnimalPuzzleheaded71 · 2026-08-16
- Gavin Baker: NVIDIA is becoming the "central bank of AI" — VibeMarketer_ · 2026-08-16