Running DeepSeek V4 Flash on 64GB Macs: 8 tok/s Decode After llama.cpp Patches
memeka · reddit · 2026-08-13
A developer successfully ran DeepSeek-V4-Flash-0731 on an M1 Max 64GB MacBook. By patching llama.cpp, using IQ3-XXS quantization (104GB), and limiting context to 64k, they achieved 8 tok/s decode and 30 tok/s prefill, performing much better than expected.
More from Infra
- Heron Power Invests Over $100M in US Factory for AI Datacenter Grid Tech — espricewright · 2026-08-13
- SK Hynix, Samsung, and Micron Reportedly Sell Out All 2027 HBM Capacity — Beth_Kindig · 2026-08-13
- NVIDIA Releases AI Tokenomics Guide: Turning Compute into Revenue — nvidia · 2026-08-13
- NVIDIA Defines the AI Era: AI Factories as New Infrastructure, Tokens as New Commodities — nvidia · 2026-08-13
- Kubernetes DRA Reaches GA: Native GPU Scheduling and Slicing — sloppenheimer · 2026-08-13
- SanDisk Targets 80% Long-Term Gross Margins and Revenue Growth Through 2030 — firstadopter · 2026-08-13