Low Effort Can Burn More Tokens: Qwen3.8-27B-pi Fine-Tunes Coding Agent Effort Ordering
lmoroney · x · 2026-10-06
Coding agent effort dials are hard to trust: a lower setting sometimes burns more reasoning tokens while solving fewer tasks.
Qwen3.8-27B-pi, a community fine-tune of Qwen3.8-27B for the Pi coding agent harness, trains the ordering directly:
- Stage one: supervised fine-tuning on successful, complete Pi sessions.
- Stage two: GRPO rewarding correctness first, then penalizing low/medium attempts that use more reasoning than a successful higher-effort attempt on the same task; xhigh gets no length pressure.
- Self-reported: medium matches the base model's xhigh completion rate with about 41% fewer output tokens, though the base keeps an edge on one chart.
Practical tip: run the same task set at each effort level and chart pass rate vs tokens — both lines should climb with effort, otherwise don't trust the dial.
More from coding & agent
- Dev Uses OpenAI Codex to Build a Two-Player Game Boy Link-Cable Paintball Game — pvncher · 2026-10-06
- Building an Interactive 3D Animal Atlas with GPT-6 Astra, No 3D Software Needed — CodeByPoonam · 2026-10-06
- Claude Opus 5.5 reportedly much better at asking permission in auto mode — stefanjblos · 2026-10-06
- Anaconda launches agent swarms and autonomous red-team agents for AI dev — anacondainc · 2026-10-06
- Asking your agent 'what are the options' unlocks the good-engineer headspace — _Stocko_ · 2026-10-06
- Forge with Rigor: specs are canonical, gate every stage, nothing certifies its own work — 53 findings at 13% token cost — colinmcnamara · 2026-10-06