Developer Runs 291B Model Locally on Four Mac Studios

eptwts · x · 2026-08-28

Developer @stevenzhang built a local inference cluster using four Mac Studios. By tensor-sharding across four Ultra chips via a Thunderbolt 5 ring with RDMA, the setup runs a 291B parameter model. This approach achieves zero inference costs without relying on cloud providers.

Original post →

More from Infra

Infra channel →