Gemma 4 26B Doubles Inference Speed on Mac
Google announced that community-driven MLX optimizations have doubled the inference speed of the Gemma 4 26B A4B model on Mac, reaching 560 tps decode and 7k tps prefill with 8 concurrent requests, a 130.7% improvement over the baseline.
2026-09-02 ~ 2026-09-02 · 3 related posts
- Gemma 4 26B A4B inference on Mac is now 2x faster — GlennCameronjr · 2026-09-02
- Gemma 4 runs 2x faster on Mac — gajesh · 2026-09-02
- Gemma 4 speed doubled on MLX, achieving 130.7% performance gain — TheMoonMidas · 2026-09-02