Solo fine-tune fixes Qwen3.8-27B: stops low-effort reasoning from burning more tokens than medium

victormustar · x · 2026-10-01

Developer bytkim released Qwen3.8-27B-pi, a Qwen3.8-27B fine-tune for coding work in the Pi agent harness. The motivation: the base model's effort levels were out of order — low often spent more reasoning tokens than medium, and on Terminal-Bench 2.1 medium passed fewer tasks than low.

Training pipeline:

Results: mean reasoning tokens and pass rates are now non-decreasing from low → medium → xhigh on Terminal-Bench 2.1, GPQA Diamond and SciCode. Pi at medium matches base at xhigh on Terminal-Bench (67/89) with 41% fewer output tokens; at xhigh on SciCode it uses 23% fewer output tokens. Weights available in BF16, FP8 (30.4 GB), and 17 GGUF variants (9–29 GB).

Related event: Dev Fine-tunes Qwen3.8-27B-pi to Fix Reasoning Effort Ordering(2 posts)→

Original post →

More from coding & agent

coding & agent channel →