Running DeepSeek v4 PRO on Mac M5 Max via SSD Streaming and DGX Station Optimization

antirez · x · 2026-08-16

The author explores using DGX Station to pool RAM and VRAM for large model inference, noting that while bandwidth is a bottleneck, good orchestration via DwarfStar allows Q2/Q3 quants to run well. A demo shows DeepSeek v4 PRO Q2 running on a 128GB Mac M5 Max via SSD streaming, successfully translating an Italian short story to English, proving the model's usability for such tasks.

Related event: DeepSeek v4 PRO Tested on Mac M5 Max via SSD Streaming(2 posts)→

Original post →

More from Infra

Infra channel →