M5 Ultra LLM test: 4x faster prompt processing, but double the power draw
DigitalguyCH · reddit · 2026-09-23
A Reddit user posted early benchmarks of running LLMs on Apple's M5 Ultra (video linked in original):
- Prompt processing: up to 4-4.5x faster than M3 Ultra depending on context length
- Generation: 1.5x faster tok/s
- Trade-offs: the machine draws twice the power (400W vs 200W), runs louder fans, and runs much hotter
Big performance gains, but energy efficiency regressed — a real trade-off for quiet, low-heat local deployment.
More from Embodied
- XPENG's IRON robot demos full-duplex speech with 9-mic array and lip reading — ChrisGPT · 2026-09-23
- Uber riders can now hail Waymo robotaxis on Austin freeways — ATTlKA · 2026-09-23
- Figure CEO: Humanoid Robotics Needs Four Stages and Eventually Hundreds of Billions — adcock_brett · 2026-09-23
- InstinctFlash: open-source runtime runs 8 robot model families in real time on one commercial GPU — chris_j_paxton · 2026-09-23
- M5 Ultra 96GB vs RTX 5090-6000 Pro: which rig for local video generation at $6.5K-$27K — MaxwellHusk · 2026-09-23
- DJI drone teardown shows 79% margins, and why the US can't build a DJI — rbhar90 · 2026-09-23