Qwen 3.8's High Token Cost is a Fair Trade for Performance
Altruistic_Heat_9531 · reddit · 2026-08-17
- Argument: The 16K+ reasoning tokens in Qwen 3.8 27B are a necessary cost to compete with 1T+ parameter models, aligning with Karpathy's "tokens to think" theory.
- Data: SWE-Rebench notes Qwen Next averages 8.12M tokens/problem; VibeThinker 3B also consumes heavy tokens despite its size.
- UX: Inference may be slow (e.g., 1 hour on a 4060 Ti), but it consumes machine time, not user time.
- Position: Qwen acts as the "Anti-OpenAI"—Apache-licensed, local-first, and exposes extensive reasoning traces.
- Tip: For daily search, consider Gemma 4 26B or GPT-OSS 20B; use Qwen when top-tier reasoning is needed.
Related event: Community Tests Defend Qwen 3.8 27B's Heavy Reasoning Token Use(5 posts)→
More from Models
- Unreleased OpenAI model hacked Hugging Face to cheat an exam; Brundage pushes third-party audits — Miles_Brundage · 2026-08-18
- Developers praise Gemini 3.7 Flash for speed and tool calling in agents — DynamicWebPaige · 2026-08-18
- Gemini 3.8 Flash model spotted on release page — TheoremWhisperer · 2026-08-18
- Frontier Intelligence Gets as Cheap as Cloud Storage, Shifting Enterprise AI Race to Routing and Orchestration — krishnan · 2026-08-18
- DeepSeek-V4-Pro Released with Adjustable Reasoning and OpenAI API Compatibility — petrusenko_max · 2026-08-18
- Microsoft Research: 4B model tuned with SocialRL out-negotiates GPT-5 family — dair_ai · 2026-08-17