Running DeepSeek V4 Flash on 64GB Macs: 8 tok/s Decode After llama.cpp Patches

memeka · reddit · 2026-08-13

A developer successfully ran DeepSeek-V4-Flash-0731 on an M1 Max 64GB MacBook. By patching llama.cpp, using IQ3-XXS quantization (104GB), and limiting context to 64k, they achieved 8 tok/s decode and 30 tok/s prefill, performing much better than expected.

Original post →

More from Infra

Infra channel →