TensorSharp hits 41 tok/s decoding DeepSeek V4.1 Flash on 8× A40

fuzhongkai · reddit · 2026-09-12

The developer of TensorSharp, an open-source LLM inference engine, shares optimizations for running DeepSeek V4.1 Flash on 8× NVIDIA A40 with layer-split execution:

| Metric | Optimized result |

|---|---|

| Q2K single-request decode | 41.1 tok/s |

| Q2K prefill (GPU-resident Engram test) | 532–541 tok/s |

| Q2K cold model load (MooseFS) | 144–155 s (2.5× faster) |

Key changes:

The author notes these are project benchmarks, not a claim of superiority over llama.cpp, since no compatible same-weight reference was available, and welcomes independent runs.

Original post →

More from Infra

Infra channel →