Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate
ChopSticksPlease · reddit · 2026-09-05
The poster credits the Qwen and Unsloth teams' Qwen3.8 27B UD Q4KXL quant: it runs with 100k context (Q8 KV cache) on a single 3090's 24GB VRAM and performs phenomenally well for agentic coding.
Their thesis: the real threat to Anthropic/OpenAI profits isn't another frontier model, but small fast local models that handle 80-90% of mundane coding work for hours at zero API cost. They also flag a market gap — Kimi-K3 is out of reach for most businesses and prosumers — and ask whether a dual/quad GPU server (96-192GB VRAM + >256GB DDR4) or a 2-4 machine DGX cluster could run MiniMax-M3 at >30tps for agentic coding. On their dual 3090 + 128GB rig, Qwen3.8-Flash-Next Q4 is fast but doesn't feel much smarter than the 27B, and their workhorse Minimax-M2.7 is too slow for coding at larger quants.
More from coding & agent
- Voice agent testing tools compared: Cekura, Cyara and TestMu solve different problems — Fishful_Revenge · 2026-09-05
- LinkedIn VP Breaks Down Hiring Assistant: The Problem Is Fragmented Recruiting Workflows — CodeByPoonam · 2026-09-05
- Agent frameworks are 10% of the problem—state hygiene is what breaks production systems — Deepfeet-09 · 2026-09-05
- Dev workflow: iterate plans with big models, then hand off execution to small reasoning models — intellectronica · 2026-09-05
- Burn Bar for Omarchy visualizes Claude/Codex token burn, quotas and GPU load locally — DanWahlin · 2026-09-05
- shadcn's agent for watching restaurant reservations cheated and falsely reported success — itsOmSarraf_ · 2026-09-05