Qwen3.8 27B xhigh burns 23.8K tokens overthinking a simple card prompt — and crushes every rival
hiImMate · reddit · 2026-08-18
A user tested Qwen3.8 27B (UDQ4XL quant) thinking-effort tiers against DS V4 Flash, ChatGPT, and Claude Opus 5, using an open-ended prompt to write an HTML card about apple benefits.
- xhigh spent 23.8K tokens thinking and produced a result far above everything else; medium thought only 794 tokens and output was mediocre.
- DS V4 Flash used 927 tokens; ChatGPT and Claude thought lightly and also trailed.
- With a detailed prompt specifying design specifics (orchard-note style, perforation, stats section), medium produced 3.3K tokens and a result close to xhigh quality.
Takeaway: xhigh can get into a long thinking match even on simple prompts but delivers notably better output; with concrete prompts, medium cuts tokens/time heavily while staying strong — good for low-tps setups.
More from Models
- User claims DeepSeek V4 Pro is mis-trained: cheaper Flash beats it — karminski3 · 2026-08-18
- Qwen3.8-27B Uncensored MLX build trends on Hugging Face for Apple Silicon — orcarouter · 2026-08-18
- Kimi K3 finds unpatched stack overflow in Go TS compiler — DanielLockyer · 2026-08-18
- Qwen3.8-27B Runs at 130 t/s Locally on Five-Year-Old Gaming GPUs — IgorCarron · 2026-08-18
- Sakana AI releases Japanese reasoning model Namazu on OpenRouter — SakanaAILabs · 2026-08-18
- Grok 4.6 ranks #2 in legal agent eval at 13× lower cost than GPT-5.6 — XFreeze · 2026-08-18