M4 Max Benchmarks: oMLX Wins at Long Context, Prefix Caching is Key
vitordeas · reddit · 2026-08-30
Benchmarked 5 runtime/quant stacks running Qwen3.8-27B (32K-256K) on an M4 Max (128GB).
Key Takeaways:
- ≤128K: MTPLX 4-bit (native MTP) is the throughput king.
- 256K: MTPLX verification collapses (7 tok/s); oMLX and mlx-dspark lead (14-15 tok/s).
- Biggest Speedup: Warm prefix cache cuts 256K prefill from 40 mins to 2 mins. oMLX supports edit-in-the-middle caching; MTPLX does not.
Tech Specs: Compares oMLX (AWQ/oQ8e), MTPLX (native MTP), and mlx-dspark (DFlash2 drafter).
More from Infra
- Own your harness, and if possible, own the model layer too — omarsar0 · 2026-08-30
- GLM 5.3 Flash Inference Extremely Slow on Apple Silicon — CentrifugalMalaise · 2026-08-30
- Tencent Hunyuan Hy4-preview runs in vLLM day 0: 770B MoE with 1M context — TencentHunyuan · 2026-08-30
- "Sharding Is the Transfer": MindLab Breaks 2M-Context VRAM Wall with 2D KV Resharding — 青稞AI · 2026-08-30
- Oberik: Open-source agent backend solves document parsing and multi-tenancy pain points — kahveciderin · 2026-08-29
- CUDA 'sticky errors' and undefined behavior explained — blelbach · 2026-08-29