Running DeepSeek v4 PRO on Mac M5 Max via SSD Streaming and DGX Station Optimization
antirez · x · 2026-08-16
The author explores using DGX Station to pool RAM and VRAM for large model inference, noting that while bandwidth is a bottleneck, good orchestration via DwarfStar allows Q2/Q3 quants to run well. A demo shows DeepSeek v4 PRO Q2 running on a 128GB Mac M5 Max via SSD streaming, successfully translating an Italian short story to English, proving the model's usability for such tasks.
Related event: DeepSeek v4 PRO Tested on Mac M5 Max via SSD Streaming(2 posts)→
More from Infra
- xllm generates an image in 0.4 seconds — warycat · 2026-08-24
- WULF CEO reveals modern AI data centers use minimal water via closed-loop systems — robleclerc · 2026-08-24
- Cursor Team Publishes 'Git at Any Scale', Advocating for Stateless Infrastructure — thesephist · 2026-08-24
- AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving — AccBalanced · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24
- Local development is fast and controllable, why rely solely on the cloud? — vboykis · 2026-08-24