TensorSharp Hits 500+ tok/s Prefill for DeepSeek V4.1 Flash on 8×A40 GPUs

fuzhongkai · reddit · 2026-09-13

The open-source TensorSharp inference engine added a dedicated DeepSeek V4.1 Flash path, reaching 533–539 tok/s prefill and 40 tok/s single-stream decode on 8× NVIDIA A40s (Q2K), and 451–492 tok/s prefill with 48.9 tok/s aggregate at concurrency 4 (Q4KM), with 65K context and F16 KV cache. Key findings: keeping 60 GiB of quantized Engram tables GPU-resident boosted prefill from 210 to 535 tok/s; consolidating backends cut decode scheduler splits from 570 to 8; async host warming reduced host MoE offload from 3 layers to 1; and routed-expert TP=8 was slower than layer split (22 vs 32 tok/s) without NVLink, as cross-NUMA communication dominated. Author invites discussion on Engram-style lookup, large-MoE placement, and batching.

Related event: TensorSharp Runs DeepSeek V4.1 Flash on 8x A40 with Strong Results(2 posts)→

Original post →

More from Infra

Infra channel →