Coding agent on one RTX 3090: Qwen3.8-27B throughput, four tasks, and a reasoning-budget failure
Adorable-Cost-3249 · reddit · 2026-10-02
The author ran Qwen3.8-27B Q4KM locally on a single RTX 3090 with llama.cpp and OpenCode, publishing full setup and benchmarks.
Performance
- 131K context fully on GPU, 22.2GB peak memory; generation slows from 36.4 tok/s at 2K input to 20.9 tok/s at 120K.
- With prefix cache reuse, first-token latency at 120K drops to 0.63s; a fresh 120K prompt takes 183s.
Four 8-minute Python coding tasks: three passed all independent checks (though later review still found Decimal rounding and concurrency-test flaws); the incremental build planner failed entirely — its 8,192-token output budget was consumed by reasoning alone, ending with length and no patch. Disabling thinking solved the same task in 5m40s with 9/10 passing — an output-budget failure, not a context limit.
Workflow takeaway: bounded tasks with explicit acceptance criteria, diff review, and independent checks.
More from coding & agent
- Dev says Claude Code with Opus 5.5 is 'a lot of fun': 'It was mostly in my head' — Angaisb_ · 2026-10-02
- NYT editor: AI excels at code for the exact reason it's mediocre at writing — dylfreed · 2026-10-02
- AI shifts game dev from writing code to testing and directing, vets say — AIandDesign · 2026-10-02
- Agent-Reach, Patchright Enhanced, Scrapling: 3 tools to let agents scrape any website — FinanceYF5 · 2026-10-02
- 3 open-source tools give AI agents web-scraping superpowers, even on API-less sites — FinanceYF5 · 2026-10-02
- Claude Opus 5.5 directed a 5.5-minute AI short film from one prompt in ComfyUI — Cheap_Credit_3957 · 2026-10-02