Qwen3.8-27B thinking mode burns 6x time for quality gains on M5 Max
DerTomsn · reddit · 2026-08-30
Benchmarks on Apple M5 Max show that Qwen3.8-27B consumes 5.5x more tokens and runs 6x longer when 'thinking' mode (xhigh) is enabled. However, the output quality sees a significant boost. Completely turning off thinking drastically reduces quality, placing it behind Qwen3.6-35B-A3B and other MoE models that offer much higher Tok/s.
More from Models
- heretic: fully automatic censorship removal for LLMs nears 29k stars — p-e-w · 2026-08-30
- Experiment: Claude Easily Assisted in Piracy and Reverse Engineering via agents.md — adonis_singh · 2026-08-30
- OpenAI dominates browser use while Claude's strength is mostly coding, exec says — bindureddy · 2026-08-30
- Model performance degrades in long context; token efficiency varies widely across labs — zakelfassi · 2026-08-30
- Claude Opus 5 Backlash: Benchmarks Soar But Daily Use Fails — gerardsans · 2026-08-30
- 'The curve of the letter b is invisible to the model' — tokenizer meme resurfaces — rickasaurus · 2026-08-30