Open-source Swift-Qwen3.8-27B cuts thinking tokens 58% for 1.95x speed at <1% accuracy loss

Secure_Recording_472 · reddit · 2026-09-14

UkisAI open-sourced Swift-Qwen3.8-27B, a post-trained Qwen 3.8 27B that penalizes tokens linked to overthinking loops (identified by clustering OOD traces across coding/language/vision/agentic domains) without force-shortening reasoning, then restores accuracy via on-policy distillation and RL. Results: -58% thinking tokens, 1.95x speedup, <1% accuracy loss. GGUF quants, community variants, and a free Nvidia-backed research API (5 RPM) are available on Hugging Face.

Related event: Swift-Qwen3.8-27B hits HF trending with 58% fewer thinking tokens(2 posts)→

Original post →

More from Models

Models channel →