Running Qwen3.8 at IQ3_S on 12GB VRAM: 20-30 tok/s decode, 131k context, sub-$2k rig

bodhi371 · reddit · 2026-10-09

A Reddit user got Qwen3.8-Flash-Next-GSQ-RCO-Abliterated (IQ3S) running on 12GB VRAM + 32GB RAM + NVMe: 20-30 tok/s decode (39-45 tok/s at Q2) and 300 to 90k tok/s prefill at 131k context, using an experimental custom fork of Strata. Total compute under $2k; author calls it "local Opus in some regards". Recipes on GitHub.

Original post →

More from Infra

Infra channel →