Qwen 27B Quant on a Single RTX 3090 Scores 87.6% on SWE-Bench via llama.cpp
Ok_Warning2146 · reddit · 2026-09-16
A developer shares full details of running unsloth's Qwen 27B IQ4NL quantized model with llama.cpp on a single RTX 3090 to benchmark SWE-Bench:
- Setup: jinja template, 229K context, q80 quantized KV cache, speculative decoding off, bf16 mmproj; sampling at temp 1.0 / topp 0.95 / topk 20.
- Results: 8 days for all 500 tests, 339 resolved — 87.5969%, remarkably strong for a quantized model on one consumer GPU.
- Issues: 9 LimitExceeded and 103 TimeoutExpired errors left 112 tests unfinished; the author asks the community for llama-server parameter tweaks to cut timeouts, improve the resolve rate, and speed things up beyond MTP.
More from coding & agent
- ModularRSI: Modular, benchmark-disjoint framework for generalizable agent harness self-improvement — IQuestLab · 2026-09-16
- Developer ditches Astra for 5.6 Sol high/xHigh setup with Luna subagent, says usage improved — rudrank · 2026-09-16
- Codex power users stuck: 7% quota left, 3-day wait, 20x plan paused — jasonkneen · 2026-09-16
- Open-source Orca runs 5 Claude Code agents in parallel, hits 60k GitHub stars — alex_verem · 2026-09-16
- claude-reflect: open-source tool turns your corrections into permanent memory for Claude Code — tom_doerr · 2026-09-16
- Free Complete Guide to Obsidian Automation released, covering AI agents on a 20,000-note vault — dSebastien · 2026-09-16