Splitting DeepSeek-V4.1-Flash across an M5 Ultra and 2× RTX PRO 6000

harrythunder · reddit · 2026-09-26

A practical local-deployment trick: DeepSeek-V4.1-Flash's prompt state is only 0.9 KB/token, so the model can be split at layer 20 across CUDA (2× RTX PRO 6000) and Metal (M5 Ultra) machines, connected over ordinary 1/10GbE networking. Detailed writeup linked — directly useful for anyone mixing Mac and PC hardware to run large models locally.

Original post →

More from Infra

Infra channel →