Gemma 4 Runs 151.4% Faster on Mac via Community MLX Inference Optimization

gajesh · x · 2026-09-05

The YukonMLX.fast team is seeking a Mac hosting provider to rent 10x M5 Max (128GB) machines for 3 months—bare metal with native MLX GPU access, sudo, and thermal telemetry—for their platform where developers optimize Apple Silicon inference (100-250% gains per model so far).

Their Gemma 4 MLX Challenge leaderboard shows Gemma 4 26B A4B now running 151.4% faster on Mac than the launch baseline, from 138 promoted submissions by 37 solvers. Scores weight prefill^0.25·decode^0.75, with decode carrying 75%; the top runs were built with Claude Opus, Gemini 3.8 Flash, and GPT-5.6, separated by under 1%.

Original post →

More from Infra

Infra channel →