MLX Fast Optimizes Gemma 4 26B: 160% Faster Decode, 42 Contributors
gajesh · x · 2026-09-10
The MLX Fast community improved Google's Gemma 4 26B on Mac by 160.4% through 180 optimizations. Under 8 concurrent requests, total decode throughput reached 654.2 tok/s, averaging 81.8 tok/s per request. The challenge, initiated by Darkbloom, aimed to optimize a non-frontier model, demonstrating community collaboration's viability.
More from Infra
- Day 249 of GPU Programming: Tracking Cerebras From CS-1 to WSE-3 Turbo-Powered CS-4 — blaizedsouza · 2026-09-10
- GPT-6 Astra pre-training cost estimated at $432M, full model $1-2B — scaling01 · 2026-09-10
- Gensyn builds IR3DE-AXL, a decentralized collective inference network with no central gateway — benfielding · 2026-09-10
- Sam Altman: average person may burn 500B tokens a month within six years — rohanpaul_ai · 2026-09-10
- AI server demand is splitting into three ownership-based markets; ~1-1.5M on-prem servers need refresh — BenBajarin · 2026-09-10
- Marvell Ramps Supply Chain for AI Scale, Analyst Flags Substrates as the Bottleneck Few Can Master — BenBajarin · 2026-09-10