omarchy-cluster runs the full 753B-param GLM-5.3 across four old Macs as one endpoint
natesiggard · x · 2026-10-08
Developer Joshua Warren released omarchy-cluster, an open-source project that lets networked machines run one local MLX model served as a single OpenAI-compatible endpoint. In the demo, a Mac Studio, three MacBook Pros and a Dell laptop on Omarchy jointly ran GLM-5.3 — 753B parameters, 216 GB of weights — as the full model, not a quantized one.
- Works across Apple Silicon Macs (macOS or Omarchy Linux), even an iPhone in the mix
- Nodes measure each other and automatically decide which machine runs which layers (pipeline parallelism)
- Slow, but proves idle hardware you already own can serve very large open-weight models
A reproducible recipe for giving old Macs a second life in local LLM inference.
More from Infra
- Chrome's new DecisionModel API reverse-engineered: prompts, limits and engine tests — dejanseo · 2026-10-08
- China's Power Glut Meets Data Centers; Immersion Cooling Traced to Bitcoin Miners — teortaxesTex · 2026-10-08
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08
- DWDM wavelength lock tightens from ±12.5GHz to ±3.5GHz, pushing optical interconnect costs — jwt0625 · 2026-10-08