Running Qwen3.8 at IQ3_S on 12GB VRAM: 20-30 tok/s decode, 131k context, sub-$2k rig
bodhi371 · reddit · 2026-10-09
A Reddit user got Qwen3.8-Flash-Next-GSQ-RCO-Abliterated (IQ3S) running on 12GB VRAM + 32GB RAM + NVMe: 20-30 tok/s decode (39-45 tok/s at Q2) and 300 to 90k tok/s prefill at 131k context, using an experimental custom fork of Strata. Total compute under $2k; author calls it "local Opus in some regards". Recipes on GitHub.
More from Infra
- Talk replay: speculative decoding with dflash/dspark speed-ups in llama.cpp — ngxson · 2026-10-10
- Dresden fab likely targets 7nm without EUV via immersion multi-patterning — pstAsiatech · 2026-10-10
- Samsung open-sources LittleBit: extreme quantization fits a 13B model in under 1GB — udmrzn · 2026-10-10
- Analyst says agentic CPU research lines up with NVIDIA Vera benchmark, flags scale-up vs scale-out question — BenBajarin · 2026-10-10
- Container cuts agent runtime P95 latency to 731ms, down ~180ms in two weeks — ritakozlov · 2026-10-10
- Amazon drops data center NDAs as community backlash spurs hundreds of moratoriums — TechCrunch AI · 2026-10-10