Swift-Qwen3.8-27B Cuts Thinking Tokens 58.3% at 1.95x Speed with Under 1% Accuracy Loss

kisjovan · hn · 2026-09-16

A team post-trained Qwen 3.8 27B to be more efficient by identifying and penalizing tokens linked to overthinking—without force-shortening reasoning—then restored accuracy via On-Policy Distillation. The open-source Swift-Qwen3.8-27B shows -58.3% thinking length, 1.95x speedup, and <1% accuracy loss.

Their thesis: reasoning length matters and should not be cut by force; only the unnecessary part should be optimized. The overthinking and anxiety-like reasoning loops they target exist even in BF16 models of this size and contribute nothing to answer quality. The approach complements, rather than replaces, reasoning effort settings, chat templates, and token caps.

The model hit 80k downloads in 3 days with independent community evals. They offer a free research-purpose OpenAI-compatible API (5RPM, GPUs courtesy of Nvidia), full GGUF Q1-Q8 quants, plus community MLX, NVFP4, W4A16, and uncensored variants.

Original post →

More from Models

Models channel →