469-question eval of Qwen3.8-27B fine-tunes: best local model 28x slower than Opus 5.5 at 99.6%

norenEnmotalen · reddit · 2026-10-08

The author ran a 469-question domain-specific eval of Qwen3.8-27B fine-tunes vs frontier models under llama.cpp with uniform thinking settings:

Takeaway: run your own domain evals; the author proposes a "PTA index" (time/tokens/accuracy) as a personal benchmarking KPI.

Original post →

More from coding & agent

coding & agent channel →