MLX Tested: M2 Max Hits 514 tok/s On-Device
MaziyarPanahi · x · 2026-07-04
To verify the "729 tok/s" claim, the author benchmarked four v2 MLX builds under identical conditions: same clinical corpus, single stream, batch 1, and decoding overhead included. Results show 8-bit outperforms fp16 on both models, hence its use in the demo. A standard M2 Max laptop achieved 514 tok/s entirely on-device.
More from Infra
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27