Two M3 Ultras linked by Thunderbolt RDMA for local tensor parallelism
antirez · x · 2026-07-20
A first test of DwarfStar distributed tensor parallelism over Thunderbolt RDMA on two M3 Ultra machines.
For DeepSeek V4 Flash Q2, the setup improved prefill throughput from 593 tok/s to 642 tok/s, while decode dropped from 39 tok/s to 33 tok/s. For Q4, prefill rose from 589 tok/s to 677 tok/s (+14.9%), while decode went from 35.5 tok/s on a single M3 Ultra to 31.4 tok/s with tensor parallelism, implying a small decode penalty for the gain in prefill speed.
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11