M4 Max Benchmarks: oMLX Wins at Long Context, Prefix Caching is Key

vitordeas · reddit · 2026-08-30

Benchmarked 5 runtime/quant stacks running Qwen3.8-27B (32K-256K) on an M4 Max (128GB).

Key Takeaways:

Tech Specs: Compares oMLX (AWQ/oQ8e), MTPLX (native MTP), and mlx-dspark (DFlash2 drafter).

Original post →

More from Infra

Infra channel →