Qwen3.8-27B thinking mode burns 6x time for quality gains on M5 Max

DerTomsn · reddit · 2026-08-30

Benchmarks on Apple M5 Max show that Qwen3.8-27B consumes 5.5x more tokens and runs 6x longer when 'thinking' mode (xhigh) is enabled. However, the output quality sees a significant boost. Completely turning off thinking drastically reduces quality, placing it behind Qwen3.6-35B-A3B and other MoE models that offer much higher Tok/s.

Original post →

More from Models

Models channel →