Benchmarks Confirm ThinkingCap & Swift Cut Qwen3.8-27B Reasoning Tokens ~40% With Minimal Loss
returnity · reddit · 2026-09-25
An independent Aider eval suite comparison of ThinkingCap-Qwen3.8-27B, Swift-Qwen3.8-27B, and vanilla Qwen3.8-27B validates both fine-tunes' claims of 40% reasoning token reduction with minimal performance loss.
Key results (2 runs per model, ±2-3% error bars):
- ThinkingCap scores identically to the original (27.1% first-try / 77.6% retry pass) while median completion tokens drop from 12,547 to 7,436, and wall time nearly halves (777s vs 1,481s per case)
- Swift posts the highest first-try pass rate (30.8%) with 7,301 median tokens
Nuances:
- ThinkingCap uses 8.5% more mean tokens due to a long tail of overthinking cases; Swift reduces more uniformly
- Both spend more tokens on failures than successes, more pronounced for Swift (13.2k vs 5.9k)
- By language: ThinkingCap trails Swift 8% on C++ but leads by 4% (JavaScript) and 6% (Python)
- ThinkingCap is the only model with a perfect well-formed-diff score
More from Models
- Blogger: Opus 5.5's strength suggests xAI's rumored Astra is smaller than believed — scaling01 · 2026-09-25
- LiquidAI extends lossless speculative decoding to vision-language models — JosephJacks_ · 2026-09-25
- GPT-5.2 solves a COLT 2022 open problem the researcher had chased since 2016 — kfountou · 2026-09-25
- Agents bypass monitoring guardrails with strategies that improve as reasoning effort scales — maksym_andr · 2026-09-25
- micro1 launches flow-transform 1.0, hits 96.0% F1 on PrivacyBench PII transformation — omarsar0 · 2026-09-25
- UkisAI ships Swift reasoning LLM family: -63.4% thinking tokens at 1.8x speed on Qwen base — Secure_Recording_472 · 2026-09-25