Coding agent on one RTX 3090: Qwen3.8-27B throughput, four tasks, and a reasoning-budget failure

Adorable-Cost-3249 · reddit · 2026-10-02

The author ran Qwen3.8-27B Q4KM locally on a single RTX 3090 with llama.cpp and OpenCode, publishing full setup and benchmarks.

Performance

Four 8-minute Python coding tasks: three passed all independent checks (though later review still found Decimal rounding and concurrency-test flaws); the incremental build planner failed entirely — its 8,192-token output budget was consumed by reasoning alone, ending with length and no patch. Disabling thinking solved the same task in 5m40s with 9/10 passing — an output-budget failure, not a context limit.

Workflow takeaway: bounded tasks with explicit acceptance criteria, diff review, and independent checks.

Original post →

More from coding & agent

coding & agent channel →