Streaming 1.6TB of K3 weights on an M5 Max 128GB is still slow
antirez · x · 2026-07-29
The post says streaming the official K3 Hugging Face weights is “a bit slow,” even when run directly on an M5 Max with 128GB of memory. The key detail is the scale of the model assets: 1.6TB of weights in mxfp4 format.
Related event: Tinkering with Local Quantized K3 Inference on Mac Hardware(2 posts)→
More from Infra
- China’s semiconductor push adds immersion DUV and near-Micron DRAM capacity — Distinct-Question-16 · 2026-07-29
- Low-VRAM LoRA training now fits DeepSeek-V4-Flash into 90 GiB VRAM — woct0rdho · 2026-07-29
- YC startup hwintelligence launches Wave, a waveform debugger with a verification agent — ycombinator · 2026-07-29
- Broadcom says AI could cut exploit windows from weeks to hours — therealdanvega · 2026-07-29
- camelAI moved its agent from VMs to a Cloudflare Durable Object — irvinebroque · 2026-07-29
- LLM Inference Handbook collects deployment, GPU, and optimization guidance for production teams — carrycooldude · 2026-07-29