DeepSeek v4 Flash GGUF quant released for DS4 engine, doubles speed to 30+ tok/s

returnity · reddit · 2026-08-01

Developer nazeshinjite released a GGUF quantized version of DeepSeek v4 Flash (DwarfStar) with a separate DSpark MTP head. Running on the DS4 engine, it achieves over 30 tok/s on MBP M5 Max, double llama.cpp's speed, making it suitable for agentic workflows. The quant uses antirez's Q2-Q4 mixed imatrix recipe, and the author seeks community feedback for performance improvements.

Original post →

More from coding & agent

coding & agent channel →