Open-source Swift-Qwen3.8-27B cuts thinking tokens 58% for 1.95x speed at <1% accuracy loss
Secure_Recording_472 · reddit · 2026-09-14
UkisAI open-sourced Swift-Qwen3.8-27B, a post-trained Qwen 3.8 27B that penalizes tokens linked to overthinking loops (identified by clustering OOD traces across coding/language/vision/agentic domains) without force-shortening reasoning, then restores accuracy via on-policy distillation and RL. Results: -58% thinking tokens, 1.95x speedup, <1% accuracy loss. GGUF quants, community variants, and a free Nvidia-backed research API (5 RPM) are available on Hugging Face.
Related event: Swift-Qwen3.8-27B hits HF trending with 58% fewer thinking tokens(2 posts)→
More from Models
- Grok trains on user data by default; business plans can opt out — carlosdponx · 2026-09-15
- Bolt Forge launches free until Oct 14 with GLM, DeepSeek and Kimi plus up to 50x more usage — HeyAmit_ · 2026-09-15
- GPT-6 Astra tested on robot control: impressive on simple tasks, limited dexterity — DJiafei · 2026-09-15
- SOTA Inference Is Nearly Free for Consumers, So the Local-Model Trend May Reverse — mobileraj · 2026-09-15
- Cursor user switches to Claude Code, burns through quota by day 3 — jdluk87 · 2026-09-15
- Claude Max and Codex tiers are creating a computing power gap that locks out $20-budget newcomers — IndraVahan · 2026-09-15