Qwen 2.5 72B local test shows 50% throughput drop at long context

julianharris · x · 2026-09-01

Tested Qwen 2.5 72B (Qwen 3.8) via MTPLX on a MacBook Pro M5 Max with 128GB RAM. Results show 27B model throughput drops significantly as the context window expands. Flash-next (Qwen 4 preview) drops much less, performing well though not matching cloud Opus.

Related event: Qwen 3.8 Tested on 128GB Mac: Smarter but Memory-Hungry(2 posts)→

Original post →

More from Models

Models channel →