1-bit Quantized Qwen 27B Runs on 12GB VRAM at 92 tokens/s

zyxciss · reddit · 2026-08-16

Reddit user zyxciss tried a 1-bit quantized version of Qwen3.8-27B (AtomicChat/Qwen3.8-27B-GGUF on HuggingFace) and managed to run it on an RTX 3060 12GB. With MTP enabled, it achieves about 92 tokens/s.

Although the model quality is severely degraded (currently unusable), the speed on a 3060 is surprising. The user had hoped for a 35B MoE version, but that seems unlikely.

Original post →

More from Infra

Infra channel →