Running DeepSeek V4 Flash on MacBook M5 Pro Hits 17 t/s via SSD Streaming
vogelvogelvogelvogel · reddit · 2026-08-06
A developer shared their experience running DeepSeek V4 Flash locally on a MacBook M5 Pro 64GB. By leveraging the SSD streaming mode of antirez's ds4 DwarfStar, non-routed weights stay resident in RAM while routed experts are pulled from the SSD on cache misses, achieving a usable generation speed of 10-17 tokens/s.
Technical Details:
- Experts and the output head are kept at Q80, while the router, embeddings, and auxiliary modules remain in FP16.
- Compiled with Metal, utilizing the --ssd-streaming and --nothink flags.
More from Infra
- RTX 3090 Test: INT8 Quantization Doubles MiniMax Video Generation Speed — Nevaditew · 2026-08-06
- Voice AI Agent Inference: Why Serverless Platforms Fall Short — Comprehensive_Quit67 · 2026-08-06
- Together AI Launches Kimi K3 API, Leads Key Inference Benchmarks — togethercompute · 2026-08-06
- RTX 3090 MiniMax H3 Benchmark: INT8 is Twice as Fast as FP8 with No Visual Loss — Nevaditew · 2026-08-06
- Post-Train Models for 10% Token Efficiency to Cut Inference Costs — ypatil125 · 2026-08-06
- Jeff Dean and Scientists Left Google Citing TPU Infrastructure Limits on Research — firstadopter · 2026-08-06