M5 Mac 128GB local LLM benchmark: Qwen3.8-flash-next hits 40-60 tok/s, Splash hits 120

surrealerthansurreal · reddit · 2026-10-06

Using a benchmark set built from a year of real coding, agentic and gameplay tasks, the author tested local models on a 128GB M5 Mac. Qwen3.8-flash-next quantizes well into 90-95GB RAM: 40 tok/s on OMLX, 60 tok/s on MTPLX with tuning; Qwen3.6 MOE on Splash runtime reaches 120 tok/s for near-parity on everything but coding. Details and an interactive chart on the blog.

Original post →

More from Infra

Infra channel →