M5 Ultra 256GB runs GLM 5.3 Flash at 68.8 tok/s with oQ4e+MTP on oMLX
cryotic · reddit · 2026-10-05
A user benchmarked GLM 5.3 Flash on an M5 Ultra 256GB machine, hitting 68.8 tok/s decode with oQ4e quantization plus MTP on oMLX 0.7.0, and 1,878 tok/s prefill — higher than other shared results, with more optimization still possible.
More from Infra
- Cloudflare ships 46 announcements in Birthday Week, bets on agentic internet economy — threepointone · 2026-10-05
- Baseten's AI-generated inference engine beats vLLM by up to 90% on decode speed — baseten · 2026-10-05
- TEE-protected KV cache helps but won't fully stop inference price manipulation — AccBalanced · 2026-10-05
- Model Swarm Protocol bets multiple local models beat one, claiming 3-5x speedup — 3x4n1m0 · 2026-10-05
- Economists debate whether OpenRouter-style Tullock contests evolve into auctions or ad-auction-like collusion — AccBalanced · 2026-10-05
- PyTorch working group standardizes hardware accelerator integration across backends — PyTorch · 2026-10-05