Tinkering with Local Quantized K3 Inference on Mac Hardware
Developers are exploring local inference for the massive K3 weights, finding that while streaming 1.6TB on an M5 Max is slow, running Q2 quantization across two 512GB Mac Studios achieves acceptable chat speeds.
2026-07-29 ~ 2026-07-29 · 2 related posts
- Streaming 1.6TB of K3 weights on an M5 Max 128GB is still slow — antirez · 2026-07-29
- Two 512GB Mac Studios can run Q2 chat, and K3 may quantize well — antirez · 2026-07-29