Two 512GB Mac Studios can run Q2 chat, and K3 may quantize well

antirez · x · 2026-07-29

The post is about dialing in an inference graph for a large model deployment. The author says Q2 across two Mac Studios with 512GB each can already run at acceptable chat speed.

They also note that K3 was trained in MXFP4, so it may quantize well, and they plan to wait for access to a second Mac Studio 512GB. In a reply, another user says they are directly streaming the official Hugging Face 1.6 TB K3 weights in mxfp4 on an M5 Max 128GB, even if it is slow. The thread is essentially a local/edge inference and quantization experiment, not a model launch.

Related event: Tinkering with Local Quantized K3 Inference on Mac Hardware(2 posts)→

Original post →

More from Infra

Infra channel →