MLX Fast Optimizes Gemma 4 26B: 160% Faster Decode, 42 Contributors

gajesh · x · 2026-09-10

The MLX Fast community improved Google's Gemma 4 26B on Mac by 160.4% through 180 optimizations. Under 8 concurrent requests, total decode throughput reached 654.2 tok/s, averaging 81.8 tok/s per request. The challenge, initiated by Darkbloom, aimed to optimize a non-frontier model, demonstrating community collaboration's viability.

Original post →

More from Infra

Infra channel →