Apple's 512GB M5 Ultra can run 93 of 97 open-weight models locally, up to ~50% faster inference
bigaiguy · x · 2026-09-04
The author argues the real story of Apple's M5 Ultra isn't the keynote's "4.3x faster AI" claim but the 512GB unified memory at 1.2TB/s bandwidth in a Mac Studio.
Key numbers:
- M3 Ultra topped out at 819GB/s; since token generation is bandwidth-bound, the jump alone yields roughly 50% faster inference before any GPU gains.
- Per canitrun's tracker estimates (not benchmarks):
- DeepSeek R1 671B @ Q4KM: 12.7 tok/s
- DeepSeek V4 Flash 284B @ Q80: 20.8 tok/s
- Llama 4 Maverick 400B @ Q80: 14.9 tok/s
- GPT-OSS 120B at full BF16: 28 tok/s
- Kimi K2.6 1T @ Q2K: 22 tok/s (fits, but Q2 is a big quality hit)
- 93 of 97 tracked open-weight models fit fully in memory at 8K context. The four that don't are the 1.6T-2.4T giants: DeepSeek V4 Pro, Kimi K3 and Qwen3.8 2.4T.
Bottom line: the 512GB config makes the Mac Studio one of the few consumer devices that can run nearly every major open-weight model locally.
More from Infra
- Microsoft names Project Zenith: dev-focused Windows shipping with AMD Ryzen AI Halo chips — tomwarren · 2026-09-04
- Reframe render engine claims ~48x speedup over Octane and Arnold after months of kernel rewrites — D3VAUX · 2026-09-04
- Local Qwen3 8B Flash on RTX 6000 Pro: 2000 tps prefill but only 40 tps decode — AppealSame4367 · 2026-09-04
- Voz hits 5-20x faster on-device speech-to-text by heavily optimizing for Apple's ANE — pcuenq · 2026-09-04
- Gemini 3.8 Flash: same token price, different cost per task — a real-world billing test — JuggernautCritical92 · 2026-09-04
- DIT.ai launches AI token exchange routing GPT, Claude, Gemini at 30-70% below list price — SucceededMind · 2026-09-04