llama.cpp patches boost DeepSeek-V4-Flash-0731 from 3.26 to 25.91 tok/s
dyn___ · x · 2026-08-04
A developer says a set of patches for DeepSeek-V4-Flash-0731 in llama.cpp improved throughput from about 3.26 tok/s to 25.91 tok/s and cut startup time from 120 seconds to 17 seconds.
The screenshots show:
- the test machine: a single RTX 3090 24GB plus older dual-socket Xeon nodes
- a step-by-step commit trail of optimizations
- the biggest gains coming from fused MoE kernels, NUMA sharding, faster process startup, expert replication, and pinned cores
The author notes there may still be more room to improve, but llama.cpp is already the bottleneck.
More from Infra
- ARC competitor says its CUDA C stack is an order of magnitude ahead — jsuarez · 2026-08-04
- A 600MB photoreal 3D Gaussian splat streams to the browser in seconds — willeastcott · 2026-08-04
- AI agent benchmark compares 5 search APIs over MCP, with Parallel fastest and Keiro cheapest — Water_Law2005 · 2026-08-04
- Dwarkesh Patel says smarter AI models could push compute prices up 10x — Dwarkesh Patel · 2026-08-04
- Inference provider swings Kimi-K3 benchmark results, with one endpoint topping CEO-Bench — AAAzzam · 2026-08-04
- Hyperscalers' AI backlog hits $2.3T, but analyst warns of circular financing — TiernanRayTech · 2026-08-04