DeepSeek V4.1 Flash Hits 40 tok/s on 8x A40 with Open-Source TensorSharp Engine
fuzhongkai · reddit · 2026-09-13
The author of open-source inference engine TensorSharp shares DeepSeek V4.1 Flash GGUF results on 8x A40 (layer split, F16 KV cache, 65K context):
- Q2K: prefill 533–539 tok/s, single-request decode 40.3–40.7 tok/s
- Q4KM: prefill 452–492 tok/s, decode 31–32.5 tok/s; 4-request aggregate throughput 48.9 tok/s
Key optimizations: Q2K keeps 60 GiB quantized Engram tables on GPU and unifies per-GPU backends, cutting decode graph splits from 570 to 8; Q4KM keeps larger tables in host memory, using auto-warming, tighter VRAM budgeting, and token-batched decode for 1.9x prefill and 2x four-request throughput, with only 1 of 40 layers keeping routed experts on CPU.
Counterintuitively, on this no-NVLink system layer split (31–32.5 tok/s) beats experimental routed-MoE tensor parallelism (21.4–22 tok/s). The author notes these are project-reported numbers; output parity and batching effects remain unverified.
More from Infra
- Positron shipped an AI chip in 15 months; Ainek rumored at 13 — DavidBennett__ · 2026-09-13
- nic_carter rebuilds Meta datacenter scale graphic from official schematics — AccBalanced · 2026-09-13
- Cambridge's David Krueger proposes systematically dismantling the AI compute supply chain — DavidSKrueger · 2026-09-13
- Asking for real quality benchmarks of Beellama's old-cache-only quantization for local coding LLMs — RadianceTower · 2026-09-13
- CME to launch cash-settled H100 and B200 GPU futures on rental-price indexes within a month — but buyers are scarce — AccBalanced · 2026-09-13
- Modular data centers shift, not solve, the labor bottleneck: module suppliers already face worker and footprint constraints — AccBalanced · 2026-09-13