Buy an M5 Max Mac Studio (128GB, €5,849) for local Qwen now, or wait for M7 Ultra in 2028?
Mxmtm · reddit · 2026-09-08
The author weighs buying a Mac Studio (M5 Max, 128GB unified memory, €5,849) now to run qwen3.8-flash-next locally at 55 tok/s (oQ4e) for private documents, notes and coding, versus waiting. Case for buying: privacy, no subscription, still a usable machine in five years. Case for waiting: RAM prices are inflated, M7 Ultra is likely 2028, per-parameter capability keeps improving, and cheap APIs leave privacy as almost the only argument for local. They also question the memory arms race — 64GB was 'plenty' a few years ago, 128GB is today's floor — and ask whether M1/M2 Ultra owners still run local models or now own an expensive web browser.
More from Infra
- Benchmarked: Apple Core AI vs MLX for on-device LLM speed on iPhone and Mac — HankYeomans · 2026-09-08
- Tiered KV-cache offloading for self-hosted LLM inference: GPU to RAM to NVMe to S3 — Responsible-You9024 · 2026-09-08
- D-Wave Finalizes Agreement with US Commerce Dept for Up to $100M in CHIPS Act Funding — ceciletamura · 2026-09-08
- Export Controls Working? H200 Sells for 280 and B300 for 450 Overseas — teortaxesTex · 2026-09-08
- Arm launches AI Portal with pre-optimized Qwen, Gemma, YOLO models accessible to coding agents via MCP — Crescitaly · 2026-09-08
- Serverless AI: why usage-based billing can cost more for static workloads — DavidLinthicum · 2026-09-08