DeepSeek V4 Flash on M2 Ultra: lossless repack to 141GiB, 25.8 t/s beats M3 Ultra

Agusx1211 · reddit · 2026-08-23

A developer built a custom llama.cpp fork for the M2 Ultra (60 cores, 192GB) that repacks DeepSeek V4 Flash losslessly to 141 GiB — smaller than the public Q4 GGUF (freeing room for context), with byte-identical output and no KV cache quantization.

Key results:

Open-sourced at llama-cpp-ds4f-m2-ultra.

Original post →

More from Infra

Infra channel →