Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate

ChopSticksPlease · reddit · 2026-09-05

The poster credits the Qwen and Unsloth teams' Qwen3.8 27B UD Q4KXL quant: it runs with 100k context (Q8 KV cache) on a single 3090's 24GB VRAM and performs phenomenally well for agentic coding.

Their thesis: the real threat to Anthropic/OpenAI profits isn't another frontier model, but small fast local models that handle 80-90% of mundane coding work for hours at zero API cost. They also flag a market gap — Kimi-K3 is out of reach for most businesses and prosumers — and ask whether a dual/quad GPU server (96-192GB VRAM + >256GB DDR4) or a 2-4 machine DGX cluster could run MiniMax-M3 at >30tps for agentic coding. On their dual 3090 + 128GB rig, Qwen3.8-Flash-Next Q4 is fast but doesn't feel much smarter than the 27B, and their workhorse Minimax-M2.7 is too slow for coding at larger quants.

Original post →

More from coding & agent

coding & agent channel →