Qwen 2.5 72B local test shows 50% throughput drop at long context
julianharris · x · 2026-09-01
Tested Qwen 2.5 72B (Qwen 3.8) via MTPLX on a MacBook Pro M5 Max with 128GB RAM. Results show 27B model throughput drops significantly as the context window expands. Flash-next (Qwen 4 preview) drops much less, performing well though not matching cloud Opus.
Related event: Qwen 3.8 Tested on 128GB Mac: Smarter but Memory-Hungry(2 posts)→
More from Models
- Hypothesis on Opus 5 leakiness: Mixed old and new training formats — Ratter · 2026-09-01
- Opus training format shift: From plain text to XML tags — Ratter · 2026-09-01
- Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored Trends on HF — DavidAU · 2026-09-01
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory — Runjia Qian · 2026-09-01
- Qwen Team Analyzes Qwen3.8-Next Architecture Design — Qwen · 2026-09-01
- Zhipu Releases INT4 & MXFP4 Versions of GLM-5.3 Flash — HaihaoShen · 2026-09-01