Two 512GB Mac Studios can run Q2 chat, and K3 may quantize well
antirez · x · 2026-07-29
The post is about dialing in an inference graph for a large model deployment. The author says Q2 across two Mac Studios with 512GB each can already run at acceptable chat speed.
They also note that K3 was trained in MXFP4, so it may quantize well, and they plan to wait for access to a second Mac Studio 512GB. In a reply, another user says they are directly streaming the official Hugging Face 1.6 TB K3 weights in mxfp4 on an M5 Max 128GB, even if it is slow. The thread is essentially a local/edge inference and quantization experiment, not a model launch.
Related event: Tinkering with Local Quantized K3 Inference on Mac Hardware(2 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23