Two M3 Ultras linked by Thunderbolt RDMA for local tensor parallelism

antirez · x · 2026-07-20

A first test of DwarfStar distributed tensor parallelism over Thunderbolt RDMA on two M3 Ultra machines.

For DeepSeek V4 Flash Q2, the setup improved prefill throughput from 593 tok/s to 642 tok/s, while decode dropped from 39 tok/s to 33 tok/s. For Q4, prefill rose from 589 tok/s to 677 tok/s (+14.9%), while decode went from 35.5 tok/s on a single M3 Ultra to 31.4 tok/s with tensor parallelism, implying a small decode penalty for the gain in prefill speed.

Original post →

More from Infra

Infra channel →