Streaming 1.6TB of K3 weights on an M5 Max 128GB is still slow
antirez · x · 2026-07-29
The post says streaming the official K3 Hugging Face weights is “a bit slow,” even when run directly on an M5 Max with 128GB of memory. The key detail is the scale of the model assets: 1.6TB of weights in mxfp4 format.
Related event: Tinkering with Local Quantized K3 Inference on Mac Hardware(2 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23