Two 512GB Mac Studios can run Q2 chat, and K3 may quantize well
antirez · x · 2026-07-29
The post is about dialing in an inference graph for a large model deployment. The author says Q2 across two Mac Studios with 512GB each can already run at acceptable chat speed.
They also note that K3 was trained in MXFP4, so it may quantize well, and they plan to wait for access to a second Mac Studio 512GB. In a reply, another user says they are directly streaming the official Hugging Face 1.6 TB K3 weights in mxfp4 on an M5 Max 128GB, even if it is slow. The thread is essentially a local/edge inference and quantization experiment, not a model launch.
Related event: Tinkering with Local Quantized K3 Inference on Mac Hardware(2 posts)→
More from Infra
- DIY local AI server uses retired NVIDIA cards for about $165 total — blelbach · 2026-07-29
- Cheap local intelligence could shift AI workloads away from the cloud — PeterDiamandis · 2026-07-29
- Bull case says AMD profit could 10x as AI spend and inference demand scale — AccBalanced · 2026-07-29
- Cradle Codec compresses KV cache for Ethernet transport between GPU nodes — knowrohit07 · 2026-07-29
- Bittensor raises q from 0.61 to 0.75, easing its emission gate for mid-ranked subnets — markjeffrey · 2026-07-29
- OpenRouter’s moat comes from routing data and tooling it can refine multiple times a day — mmurph · 2026-07-29