M3 Max 36GB local LLM shootout: Qwen3.6 MoE hits 110 tok/s on Rapid-MLX
Hyungsun · reddit · 2026-10-05
The author bought a 14-core M3 Max MacBook Pro (36GB) for $2,498 and benchmarked four local inference engines — oMLX 0.7.0, Rapid-MLX 0.15.5, Splash 1.2.1, MTPLX 2.12.2 — at temperature 0, three runs per config, medians reported.
Key decode results (natural prose prompt):
- Qwen3.6-35B-A3B (MoE): Rapid-MLX fastest at 110.2 tok/s, oMLX close at 104.2, Splash/MTPLX 77-79
- Qwen3.8-27B (dense): Splash 31.7 and MTPLX 31.5 effectively tied; oMLX only 17.9 — right at the theoretical 18 tok/s ceiling for 16GB of weights over 300GB/s bandwidth, which the others beat via speculative decoding
- Splash is erratic on the MoE: 238 tok/s on filler text at 1.5K context, dropping to 101 at 5.5K; Rapid stays 118
Other findings:
- First token on a 5.5K prompt: 4-5s on the MoE, 34s on dense (prefill 1.2k vs 150 tok/s)
- Fans stay loud under sustained inference
- The 125B Qwen3.8-Flash-Next needs 96GB+ and won't run on 36GB
- Engines persist KV caches across restarts, which can fool double benchmarks
Bottom line: pick the engine that's fastest for the model you use most; margins are otherwise small.
More from Infra
- Relace AI Reportedly Served 2.3T Tokens on OpenRouter in a Day, Topping OpenAI's 1.8T — stuffyokodraws · 2026-10-06
- Cloudflare adds natural-language observability charts, plus 9 ready-to-use debugging prompts — hichaelmart · 2026-10-06
- Cloudflare's X402 protocol makes AI agents pay per scrape, feeding an agentic economy on stablecoins — kleffew94 · 2026-10-06
- YuE2 music model runs fully in-browser: 40s song in ~60s on a MacBook via WebGPU — realmrfakename · 2026-10-06
- DeepSeek-style agent infrastructure ran 1.3B sandboxes in 4 weeks, peaking at 170K concurrent — teortaxesTex · 2026-10-06
- Perplexity CEO: power users now burn $10k+/month of compute in agent loops — rohanpaul_ai · 2026-10-06