M5 Mac 128GB local LLM benchmark: Qwen3.8-flash-next hits 40-60 tok/s, Splash hits 120
surrealerthansurreal · reddit · 2026-10-06
Using a benchmark set built from a year of real coding, agentic and gameplay tasks, the author tested local models on a 128GB M5 Mac. Qwen3.8-flash-next quantizes well into 90-95GB RAM: 40 tok/s on OMLX, 60 tok/s on MTPLX with tuning; Qwen3.6 MOE on Splash runtime reaches 120 tok/s for near-parity on everything but coding. Details and an interactive chart on the blog.
More from Infra
- Ex-UK energy official warns AI needs 500GW by 2035 and the industry isn't ready — ShakeelHashim · 2026-10-06
- Starlink now has 11,000+ satellites in orbit, two-thirds of all active satellites — XFreeze · 2026-10-06
- Ben Bajarin: Agentic AI Will Drive Datacenter CPU Demand, Scale-Up Domain Is the Key Battleground — BenBajarin · 2026-10-06
- llama.cpp adds DFlash speculative decoding for Qwen3.8-27B, faster than MTP — victormustar · 2026-10-06
- GLM 5.3 full NVFP4 deployable on 4x B200 or H200 with Marlin kernels — TheZachMueller · 2026-10-06
- Dev quantizes GLM-5.3-UNCENSORED to MXFP4, cutting size 44% for AMD GPUs — bakatristan · 2026-10-06