M5 Ultra 80-core tested with GLM-5.3-Flash: RAM is great, GPU is the bottleneck
dreamingwell · reddit · 2026-09-25
A user shares multi-round agentic inference benchmarks on the 256GB, 80-core M5 Ultra Mac Studio, reporting satisfying results running GLM-5.3-Flash at high local speeds.
Key takeaways:
- The large unified memory is clearly valuable for running big models
- The GPU is underpowered relative to the RAM and becomes the bottleneck
- He questions whether a 512GB unit makes sense for AI inference at all, since the GPU would remain the limiting factor
A useful first-hand reference for anyone considering a high-RAM Mac for local inference.
More from Infra
- Musk details xAI compute: Colossus 2 to hit 880k GB300s by year-end — elonmusk · 2026-09-25
- Qwen-Image-2.1 gets GGUF quantization, could run text-to-image on a Snapdragon 865 phone — ResidentAping · 2026-09-25
- Deep Inference-Query Engine Integration: Custom Scheduler and Workload-Aware KV Cache for Prefill-Only AI Filters — charles_irl · 2026-09-25
- Finance worker seeks local AI setups to cut soaring Codex/ChatGPT costs — Startup__Sam · 2026-09-25
- Diffusion LLM goes production: Augment Code's Mercury 2.5 switch cuts latency 82%, cost 90% — cen6wkf · 2026-09-25
- Oracle says force majeure notice doesn't signal data center delays, rent still due — AIFlow_ML · 2026-09-25