Solo fine-tune fixes Qwen3.8-27B: stops low-effort reasoning from burning more tokens than medium
victormustar · x · 2026-10-01
Developer bytkim released Qwen3.8-27B-pi, a Qwen3.8-27B fine-tune for coding work in the Pi agent harness. The motivation: the base model's effort levels were out of order — low often spent more reasoning tokens than medium, and on Terminal-Bench 2.1 medium passed fewer tasks than low.
Training pipeline:
- SFT on curated successful Pi sessions to teach the full read/edit/bash coding loop;
- GRPO-based RL with a success-conditioned reward: solve the task, and lower effort levels must not reason longer than higher ones.
Results: mean reasoning tokens and pass rates are now non-decreasing from low → medium → xhigh on Terminal-Bench 2.1, GPQA Diamond and SciCode. Pi at medium matches base at xhigh on Terminal-Bench (67/89) with 41% fewer output tokens; at xhigh on SciCode it uses 23% fewer output tokens. Weights available in BF16, FP8 (30.4 GB), and 17 GGUF variants (9–29 GB).
Related event: Dev Fine-tunes Qwen3.8-27B-pi to Fix Reasoning Effort Ordering(2 posts)→
More from coding & agent
- Early look: Codex Plugin Extensions are cool but still rough around the edges — Angaisb_ · 2026-10-01
- Retriever AI's self-run benchmark scores 24/24 vs OpenAI Dots' 22/24 on five everyday agent tasks — quarkcarbon · 2026-10-01
- Dev builds agent testing tool that catches fake tool calls and loops in full conversations — mbtigeekjung · 2026-10-01
- Dev to agent builders: stop ad channels, kill dark-pattern subscriptions for us — SuB8u · 2026-10-01
- $3,000/month multi-model workflow claims 10-50x productivity gains—and white-collar doom — opmgyhx · 2026-10-01
- "I told my AI tool to replace itself with a cheaper one" — dev celebrates friction-free agent swaps — yacineMTB · 2026-10-01