469-question eval of Qwen3.8-27B fine-tunes: best local model 28x slower than Opus 5.5 at 99.6%
norenEnmotalen · reddit · 2026-10-08
The author ran a 469-question domain-specific eval of Qwen3.8-27B fine-tunes vs frontier models under llama.cpp with uniform thinking settings:
- Opus 5.5 topped the list at 99.6% accuracy; no model hit 100%. Best local pick: mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4KM — a Q4KM quant outperforming L/XL variants.
- Speed gap is massive: 9 hours locally vs 19 minutes on frontier (28x slower), in exchange for 100% privacy vs 0 on API providers.
- Astra shows the lowest token usage but its reasoning appears masked by the provider; author calls on OpenRouter to mandate quantization disclosure.
- gpt-oss-120b badly failed pandas/numpy tasks — only domain-specific evals surface such blind spots.
- Eval tool open-sourced at github.com/ashe-wb/tuieval.
Takeaway: run your own domain evals; the author proposes a "PTA index" (time/tokens/accuracy) as a personal benchmarking KPI.
More from coding & agent
- Agent demo wires up its own WorkOS account and full SaaS stack in under 5 minutes — jeff_weinstein · 2026-10-08
- 6-day, 31-PR chat compacted 27 times still works: dev argues context compacting is basically solved — holdenmatt · 2026-10-08
- Kare: an on-prem GitHub Copilot gateway on Arduino VENTUNO Q with dynamic Qwen/Copilot/Foundry routing — unixterminal · 2026-10-08
- Evolvent AI releases RSIGym and RSI-Index: benchmarking AI self-improvement, Opus 5 leads at 0.4809 — cihangxie · 2026-10-08
- AGI House and Coframe host a build day for the self-adapting 'living internet' — agihouse_org · 2026-10-08
- Solo dev builds three.js desert racer with Opus 5.5: a 4-day AI-assisted game dev log — Promptmethus · 2026-10-08