Gemma 4 speed doubled on MLX, achieving 130.7% performance gain
TheMoonMidas · x · 2026-09-02
Gemma 4 optimization on the MLX framework has reached 560 tps decode and 7k tps prefill across 8 concurrent requests, representing a 130.7% performance gain over the baseline. These improvements are set to be integrated into @darkbloomai.
Related event: Gemma 4 26B Doubles Inference Speed on Mac(3 posts)→
More from Infra
- Deploying DeepSeek-V4 on Blackwell: Fixing 3 critical SGLang bugs — shrug_hellifino · 2026-09-02
- LangChain fine-tunes Qwen for agent evals, cutting costs by 100x vs GPT-5.5 — LangChain · 2026-09-02
- Dell's $95B AI server backlog is profitable, defying zero-margin fears — TiernanRayTech · 2026-09-02
- Anthropic Changes Messages API to Protect Against Distillation — maksym_andr · 2026-09-02
- B300 Supports 6x More Concurrent Agents Than H200: Benchmark — ryanshrout · 2026-09-02
- AMD's Primus tuning agent predicts config performance to save thousands of GPU-hours — PyTorch · 2026-09-02