Solo fine-tune of Qwen3.8-27B-pi fixes effort ordering, saves 41% tokens
victormustar · x · 2026-10-01
A solo developer fine-tuned Qwen3.8-27B into pi on rented GPUs for the Pi coding agent, fixing the base model's broken effort ordering (low often reasoning longer than medium). Two-stage recipe: SFT on curated successful Pi sessions, then GRPO RL with a success-conditioned reward enforcing effort ordering. Result: non-decreasing reasoning tokens and pass rates from low→xhigh on Terminal-Bench 2.1, GPQA Diamond and SciCode; pi at medium matches base at xhigh (67/89) with 41% fewer output tokens, and 23% fewer on SciCode xhigh. BF16, FP8 (30.4 GB) and 17 GGUF variants released.
Related event: Dev Fine-tunes Qwen3.8-27B-pi to Fix Reasoning Effort Ordering(2 posts)→
More from coding & agent
- Dots now tap your Codex and ChatGPT context and can run multiple tasks at once — dkundel · 2026-10-01
- Mathematician details agentic math workflow with Codex CLI and Claude Code, results coming — burny_tech · 2026-10-01
- Fathom: An Individuality System for Agents That Matches Top Memory Systems — allisonmaybe · 2026-10-01
- Airbench Crowdsources a Local LLM Leaderboard via One-Prompt Agent Benchmarks — dh7net · 2026-10-01
- Vercel Ship SF Agenda: Notion's Agentic Platform, Grok Powering 300K Apps, Vercel's Eve Agent Framework — evilrabbit_ · 2026-10-01
- Azure AI kicks off series: why content extraction matters more as GenAI models get stronger — adnan_hashmi · 2026-10-01