Swift-Qwen3.8-27B Cuts Reasoning Tokens 40%: Aider Benchmarks Show 2x Speed at Same Accuracy
returnity · reddit · 2026-09-16
UkisAI released Swift-Qwen3.8-27B, targeting Qwen3.8-27B's overthinking: the team identified reasoning-marker tokens that trigger overthinking in Qwen's reasoning rollouts, penalized them with RL, and added a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B (a fine-tune already proven to save 40% tokens at comparable quality).
The poster independently verified the claims with an Aider-based coding eval suite: Pass1 30.8% vs 27.1%, Pass2 75.7% vs 77.6% — essentially parity — while completion tokens dropped from 12,547 to 7,301 (63%) and seconds per case halved from 1,481 to 750. Because decode speed degrades with longer responses, token savings translate into outsized time savings. A big win for local 27B users.
More from Models
- OpenAI to host GPT-6 Community Night in London, confirming 'GPT-6 Astra' codename — paw_lean · 2026-09-16
- Testing AI Work Assistants on Trip Planning: ChatGPT and Claude Both Fumble Hotel Availability — giffmana · 2026-09-16
- Dev complains frontier models remain bad at Spanish despite vendors' broken fix promises — Angaisb_ · 2026-09-16
- Expert re-grading shows physics benchmarks are broken: most 'model errors' are benchmark or grading errors — zainhas · 2026-09-16
- Most benchmark 'model errors' are actually benchmark or grading errors, analysis finds — zainhas · 2026-09-16
- Leaker Spots Unannounced 'Speech to Speech Index' Page — testingcatalog · 2026-09-16