Gemma 4 speed doubled on MLX, achieving 130.7% performance gain

TheMoonMidas · x · 2026-09-02

Gemma 4 optimization on the MLX framework has reached 560 tps decode and 7k tps prefill across 8 concurrent requests, representing a 130.7% performance gain over the baseline. These improvements are set to be integrated into @darkbloomai.

Related event: Gemma 4 26B Doubles Inference Speed on Mac(3 posts)→

Original post →

More from Infra

Infra channel →