DeepSeek V4.1 Flash Hits 40 tok/s on 8x A40 with Open-Source TensorSharp Engine

fuzhongkai · reddit · 2026-09-13

The author of open-source inference engine TensorSharp shares DeepSeek V4.1 Flash GGUF results on 8x A40 (layer split, F16 KV cache, 65K context):

Key optimizations: Q2K keeps 60 GiB quantized Engram tables on GPU and unifies per-GPU backends, cutting decode graph splits from 570 to 8; Q4KM keeps larger tables in host memory, using auto-warming, tighter VRAM budgeting, and token-batched decode for 1.9x prefill and 2x four-request throughput, with only 1 of 40 layers keeping routed experts on CPU.

Counterintuitively, on this no-NVLink system layer split (31–32.5 tok/s) beats experimental routed-MoE tensor parallelism (21.4–22 tok/s). The author notes these are project-reported numbers; output parity and batching effects remain unverified.

Original post →

More from Infra

Infra channel →