M5 Ultra 256GB vs dual DGX Spark: local LLM buyer weighs throughput vs memory
MasterNomie · reddit · 2026-10-11
A Reddit user with an RTX 5090 is torn between an M5 Ultra 256GB and dual DGX Spark for local LLM inference.
- Their research so far: token generation is close or slightly favors the M5 Ultra, while dual Sparks lead on prompt processing and parallel/multi-user requests
- Dual Spark suits agentic coding with many parallel requests, at the cost of roughly 1/3 of the 5090's decode speed
- The upside is running much larger models at the 256GB unified-memory tier
- The post cites a Mindstudio comparison and asks actual owners for prompt processing/decode benchmarks
More from Infra
- Anthropic's Massive Queensland Data Centre Draws Local Attention — Whitehatnetizen · 2026-10-11
- 5 LLM deployment patterns every AI engineer should know, from API calls to hybrid — goyalshaliniuk · 2026-10-11
- Neoclouds that just buy power and GPUs will lose to software-first players, says Beam founder — edgarpavlovsky · 2026-10-11
- AI Infrastructure Borrowing Slumps 80% in Three Months, From $113B to $23B — TansuYegen · 2026-10-11
- OpenAI, Microsoft, Google and 4 others signed a pledge to cover AI data centers' grid costs — ChrisGPT · 2026-10-11
- One-town monopolies: Yiwu makes 80% of Christmas decor, every EUV machine comes from Veldhoven — sahilypatel · 2026-10-11