MLX Tested: M2 Max Hits 514 tok/s On-Device

MaziyarPanahi · x · 2026-07-04

To verify the "729 tok/s" claim, the author benchmarked four v2 MLX builds under identical conditions: same clinical corpus, single stream, batch 1, and decoding overhead included. Results show 8-bit outperforms fp16 on both models, hence its use in the demo. A standard M2 Max laptop achieved 514 tok/s entirely on-device.

Original post →

More from Infra

Infra channel →