llama.cpp patches boost DeepSeek-V4-Flash-0731 from 3.26 to 25.91 tok/s

dyn___ · x · 2026-08-04

A developer says a set of patches for DeepSeek-V4-Flash-0731 in llama.cpp improved throughput from about 3.26 tok/s to 25.91 tok/s and cut startup time from 120 seconds to 17 seconds.

The screenshots show:

The author notes there may still be more room to improve, but llama.cpp is already the bottleneck.

Original post →

More from Infra

Infra channel →