实测证实:两款微调版 Qwen3.8-27B 推理 token 减约 40% 几乎不掉分
returnity · reddit · 2026-09-25
An independent Aider eval suite comparison of ThinkingCap-Qwen3.8-27B, Swift-Qwen3.8-27B, and vanilla Qwen3.8-27B validates both fine-tunes' claims of 40% reasoning token reduction with minimal performance loss.
Key results (2 runs per model, ±2-3% error bars):
- ThinkingCap scores identically to the original (27.1% first-try / 77.6% retry pass) while median completion tokens drop from 12,547 to 7,436, and wall time nearly halves (777s vs 1,481s per case)
- Swift posts the highest first-try pass rate (30.8%) with 7,301 median tokens
Nuances:
- ThinkingCap uses 8.5% more mean tokens due to a long tail of overthinking cases; Swift reduces more uniformly
- Both spend more tokens on failures than successes, more pronounced for Swift (13.2k vs 5.9k)
- By language: ThinkingCap trails Swift 8% on C++ but leads by 4% (JavaScript) and 6% (Python)
- ThinkingCap is the only model with a perfect well-formed-diff score
「模型」频道最新
- 博主称 Opus 5.5 惊艳表现暗示 xAI 的 Astra 体量更小 — scaling01 · 2026-09-25
- LiquidAI 把投机解码扩展到视觉语言模型,文本图像统一处理 — JosephJacks_ · 2026-09-25
- 研究者:GPT-5.2 解决了我搁置十年的 COLT 开放问题 — kfountou · 2026-09-25
- 智能体绕过监控护栏策略曝光:推理越强越会规避监管 — maksym_andr · 2026-09-25
- micro1 发布 PII 转换模型 flow-transform 1.0,PrivacyBench 得分 96.0% F1 — omarsar0 · 2026-09-25
- UkisAI 发布 Swift 高效推理模型家族:思考 token 减少 63.4%,速度提升 1.8 倍 — Secure_Recording_472 · 2026-09-25