12GB VRAM Test: Qwen3.8 Q2 Usable, but MoE Remains Superior
CoffeeToCode99 · reddit · 2026-08-17
Author compared Qwen3.8-27B (Dense, Q2/Q3) vs. Qwen3.6-35B-A3B (MoE, Q4) on an RTX 5070 Ti Laptop (12GB VRAM), concluding:
Setup
- Hardware: RTX 5070 Ti Laptop, 12GB VRAM.
- Backend: llama.cpp CUDA.
- Settings: 4k context, q8 KV, --fit on, no MTP.
Key Findings
- Qwen3.8 Q2 (Dense)
- Faster speeds (Prompt 412 t/s, Gen 35.9 t/s).
- Quality dropped in sanity tests (e.g., failed bat-and-ball), proving Q2 has noticeable degradation.
- Qwen3.8 Q3 (Dense)
- Maintained accuracy (6/6) in sanity tests.
- Extremely slow generation (7.5 t/s), making it painful for interactive use.
- Qwen3.6 35B-A3B (MoE)
- Perfect accuracy (6/6).
- Fastest generation (59 t/s), significantly outpacing Dense Q3.
Conclusion
- If the goal is "can it run," Qwen3.8 Q2 works.
- For actual daily chat/coding, the 35B-A3B MoE remains the best choice on 12GB VRAM, balancing speed and quality.
More from Infra
- Bitcoin's 500x Supercompute Edge Powers Bittensor's Decentralized Inference — markjeffrey · 2026-08-17
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17
- CoreWeave Revenue Outpaces Big Cloud Early Stages Amid AI Pivot — a16z · 2026-08-17