mlx-dspark reportedly speeds up Gemma 4 by about 3× on M4 Pro Macs
GlennCameronjr · x · 2026-07-26
mlx-dspark reportedly makes Gemma 4 run about 3× faster on Mac in warm benchmarks.
- The screenshot shows tests on an M4 Pro using mlx-dspark with --trials 3.
- Gemma-4 12B (8-bit) reaches the largest gains: 3.09× math, 2.63× chat, 2.61× code, at roughly 49 tok/s.
- Other supported models in the table include Qwen3.6-27B, Qwen3-14B, Qwen3-8B, Qwen3-4B, Ornith-1.0-9B, and Ternary-Bonsai-27B.
- The takeaway is simple: local model serving on Mac can get a meaningful speed boost, with the biggest gains varying by model and task.
More from Infra
- How to Network Multiple PCs for Local LLM Inference: A Hardware Setup Guide — BinaryGrind · 2026-07-26
- ASML's EUV Machines: 20G Acceleration with 5-Atom Precision — ZeroStateReflex · 2026-07-26
- Jack Dorsey’s Buzz pitch centers on shared compute for open-model communities — MarvinTBaumann · 2026-07-26
- AWS, Google, Azure and Cloudflare are all building agent sandboxes — krishnan · 2026-07-26
- AI stocks slide as Alphabet lifts 2026 CapEx to $195B–$205B — tengyanAI · 2026-07-26
- MCP release candidate adds stateless scaling and enterprise auth for agents — davemccollough · 2026-07-26