Gemma 4 Runs 151.4% Faster on Mac via Community MLX Inference Optimization
gajesh · x · 2026-09-05
The YukonMLX.fast team is seeking a Mac hosting provider to rent 10x M5 Max (128GB) machines for 3 months—bare metal with native MLX GPU access, sudo, and thermal telemetry—for their platform where developers optimize Apple Silicon inference (100-250% gains per model so far).
Their Gemma 4 MLX Challenge leaderboard shows Gemma 4 26B A4B now running 151.4% faster on Mac than the launch baseline, from 138 promoted submissions by 37 solvers. Scores weight prefill^0.25·decode^0.75, with decode carrying 75%; the top runs were built with Claude Opus, Gemini 3.8 Flash, and GPT-5.6, separated by under 1%.
More from Infra
- Extropic unveils Z1T sparse models claiming up to 140x energy efficiency gains over GPUs — beffjezos · 2026-09-05
- DeepSeek to deploy at least 160,000 next-gen Huawei AI chips at massive Inner Mongolia data center — Polymarket · 2026-09-05
- Nvidia guides FY28 to ~$691B: non-hyperscaler AI customers now half of business, growing 100% a year — Beth_Kindig · 2026-09-05
- Coatue in talks to form multibillion-dollar JV with chip startup MatX to finance die purchases and foundry capacity — steph_palazzolo · 2026-09-05
- Running an Opus-level coding agent locally at 2x speed for free: a 15-page report — julianharris · 2026-09-05
- There's no agreed way to value a GPU running inference—and compute futures now settle on these indexes — AccBalanced · 2026-09-05