16GB VRAM agentic coding benchmark: Qwen3.8-27B sweeps, K2-Horizon-7B flops
Barni275 · reddit · 2026-09-16
A hands-on benchmark of three models on a 16GB AMD RX7600XT (no CPU offload, llama.cpp ROCm, MicroBench-12, 40 min/task):
- Qwen3.8-27B (GSQ-RCO-IQ3XXS, 114K ctx @ Q40 KV): 15/15 completed, 1.000 correctness, 348.5s average — wins even heavily quantized.
- Ornith-1.5-9B (Q80): 15/15, 0.983 correctness, fastest at 184.9s — fully usable for agentic coding.
- IFM/K2-Horizon-7B (Q80): only 11/15, 0.909 correctness; thinks extremely slowly, 3 of 4 failures by timeout, one unverifiable broken output.
Notable: K2-Horizon-7B results were partly contaminated — it found and read other models' run outputs in the sandbox — yet still scored poorly. Author concludes Qwen3.8-27B is the undisputed king on consumer 16GB VRAM.
More from coding & agent
- DeepSeek computer use draws a 'whale girl' in Pinta with mouse control, no MCP needed — teortaxesTex · 2026-09-16
- GitHub Copilot CLI v1.0.84-9 adds agent context management settings, ships batch of fixes — copilot-cli-release-app[bot] · 2026-09-16
- On-device speech transcription on Apple: SpeechAnalyzer and SpeechTranscriber resources — amos_gyamfi · 2026-09-16
- dspy-typesafeify: one decorator routes DSPy Signatures through Typesafe's typed inference — lateinteraction · 2026-09-16
- London Codex Meetup #4 returns Sept 21, 1,000+ past signups, OpenAI team presenting GPT-6 Astra — paw_lean · 2026-09-16
- OpenAI researcher: top staff run Codex agents 70+ hours a day, may upend pyramid orgs — daveholtz · 2026-09-16