Splitting DeepSeek-V4.1-Flash across an M5 Ultra and 2× RTX PRO 6000
harrythunder · reddit · 2026-09-26
A practical local-deployment trick: DeepSeek-V4.1-Flash's prompt state is only 0.9 KB/token, so the model can be split at layer 20 across CUDA (2× RTX PRO 6000) and Metal (M5 Ultra) machines, connected over ordinary 1/10GbE networking. Detailed writeup linked — directly useful for anyone mixing Mac and PC hardware to run large models locally.
More from Infra
- Google Drops Free GPU Masterclass as Part of DeepMind's 'How to Scale Your Model' Book — mdancho84 · 2026-09-26
- Google's $15B India AI datacenter draws farmer protests over confiscated land — nordicinst · 2026-09-26
- Nvidia's SoL-Pi cuts coding agent token usage by up to 49% by optimizing the harness — The Decoder · 2026-09-26
- Running Qwen3.8 Locally on 6x3090s with exllamav3 Hits 80-120 tok/s — takoulseum · 2026-09-26
- Gemma 4 31B UD-Q8_K_XL Suddenly Hits "Exceeds Shared Memory" at 68k Context — Few_Professional6859 · 2026-09-26
- Meta Open-Sources Muse Glimmer: 30B Agentic Model That Fits a 24GB Consumer GPU — bibryam · 2026-09-26