Solo fine-tune of Qwen3.8-27B-pi fixes effort ordering, saves 41% tokens

victormustar · x · 2026-10-01

A solo developer fine-tuned Qwen3.8-27B into pi on rented GPUs for the Pi coding agent, fixing the base model's broken effort ordering (low often reasoning longer than medium). Two-stage recipe: SFT on curated successful Pi sessions, then GRPO RL with a success-conditioned reward enforcing effort ordering. Result: non-decreasing reasoning tokens and pass rates from low→xhigh on Terminal-Bench 2.1, GPQA Diamond and SciCode; pi at medium matches base at xhigh (67/89) with 41% fewer output tokens, and 23% fewer on SciCode xhigh. BF16, FP8 (30.4 GB) and 17 GGUF variants released.

Related event: Dev Fine-tunes Qwen3.8-27B-pi to Fix Reasoning Effort Ordering(2 posts)→

Original post →

More from coding & agent

coding & agent channel →