Gemma 4 26B Doubles Inference Speed on Mac

Google announced that community-driven MLX optimizations have doubled the inference speed of the Gemma 4 26B A4B model on Mac, reaching 560 tps decode and 7k tps prefill with 8 concurrent requests, a 130.7% improvement over the baseline.

2026-09-02 ~ 2026-09-02 · 3 related posts