M5 Ultra 80-core tested with GLM-5.3-Flash: RAM is great, GPU is the bottleneck

dreamingwell · reddit · 2026-09-25

A user shares multi-round agentic inference benchmarks on the 256GB, 80-core M5 Ultra Mac Studio, reporting satisfying results running GLM-5.3-Flash at high local speeds.

Key takeaways:

A useful first-hand reference for anyone considering a high-RAM Mac for local inference.

Original post →

More from Infra

Infra channel →