M3 Max 36GB local LLM shootout: Qwen3.6 MoE hits 110 tok/s on Rapid-MLX

Hyungsun · reddit · 2026-10-05

The author bought a 14-core M3 Max MacBook Pro (36GB) for $2,498 and benchmarked four local inference engines — oMLX 0.7.0, Rapid-MLX 0.15.5, Splash 1.2.1, MTPLX 2.12.2 — at temperature 0, three runs per config, medians reported.

Key decode results (natural prose prompt):

Other findings:

Bottom line: pick the engine that's fastest for the model you use most; margins are otherwise small.

Original post →

More from Infra

Infra channel →