双 M3 Ultra 做张量并行初测
antirez · x · 2026-07-20
DwarfStar 做了一个在两台 M3 Ultra 之间通过 Thunderbolt RDMA 跑 tensor parallel 的初测。
- DeepSeek V4 Flash Q2:单机 M3 Ultra 的 prefill 为 593 tok/s,decode 为 39 tok/s;双机 tensor parallel 后,prefill 提升到 642 tok/s,decode 降到 33 tok/s。
- DeepSeek V4 Flash Q4:prefill 从 589 tok/s 提升到 677 tok/s,增幅约 14.9%;decode 则从单机 35.5 tok/s 降到 31.4 tok/s,说明这种并行方式能换来更高 prefill,但会带来一定 decode 代价。
「Infra」频道最新
- 推理服务关键指标盘点:流式速度 TPOT 与可用性如何影响体验 — abhijithneil · 2026-09-11
- PlanetScale 推出分片 Postgres,768 台服务器组成 1PB 单库 — dhruv2038 · 2026-09-11
- 7900 XTX 24GB 能否跑 Qwen,Reddit 征集 ROCm 实测数据 — thenomadexplorerlife · 2026-09-11
- 实测戳破 RTK 省钱神话:token 省了,账单分文未降 — Bartaseth · 2026-09-11
- SF Compute 创始人:现在买算力是「绝对糟糕的体验」 — IgorCarron · 2026-09-11
- 开源 SmolVM:2026 年 Agent 不止用浏览器,而是拥有整台电脑 — aniketmaurya · 2026-09-11