Swift-Qwen3.8-27B Cuts Reasoning Tokens 40%: Aider Benchmarks Show 2x Speed at Same Accuracy

returnity · reddit · 2026-09-16

UkisAI released Swift-Qwen3.8-27B, targeting Qwen3.8-27B's overthinking: the team identified reasoning-marker tokens that trigger overthinking in Qwen's reasoning rollouts, penalized them with RL, and added a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B (a fine-tune already proven to save 40% tokens at comparable quality).

The poster independently verified the claims with an Aider-based coding eval suite: Pass1 30.8% vs 27.1%, Pass2 75.7% vs 77.6% — essentially parity — while completion tokens dropped from 12,547 to 7,301 (63%) and seconds per case halved from 1,481 to 750. Because decode speed degrades with longer responses, token savings translate into outsized time savings. A big win for local 27B users.

Original post →

More from Models

Models channel →