TensorSharp Hits 500+ tok/s Prefill for DeepSeek V4.1 Flash on 8×A40 GPUs
fuzhongkai · reddit · 2026-09-13
The open-source TensorSharp inference engine added a dedicated DeepSeek V4.1 Flash path, reaching 533–539 tok/s prefill and 40 tok/s single-stream decode on 8× NVIDIA A40s (Q2K), and 451–492 tok/s prefill with 48.9 tok/s aggregate at concurrency 4 (Q4KM), with 65K context and F16 KV cache. Key findings: keeping 60 GiB of quantized Engram tables GPU-resident boosted prefill from 210 to 535 tok/s; consolidating backends cut decode scheduler splits from 570 to 8; async host warming reduced host MoE offload from 3 layers to 1; and routed-expert TP=8 was slower than layer split (22 vs 32 tok/s) without NVLink, as cross-NUMA communication dominated. Author invites discussion on Engram-style lookup, large-MoE placement, and batching.
Related event: TensorSharp Runs DeepSeek V4.1 Flash on 8x A40 with Strong Results(2 posts)→
More from Infra
- agi-memory: SQLite-only persistent memory MCP server for coding assistants, 32MB RAM — Rude_Gate7599 · 2026-09-14
- $3000 home server with 128GB VRAM runs Qwen3.8-next at 1.3k tps prefill, 70 tps code — Thin_Pollution8843 · 2026-09-14
- Top 10 foundry revenue hits $53.5B in Q2, TSMC holds 72.5% share — Beth_Kindig · 2026-09-14
- Qwen3.8 27B INT4 With 144K Context Runs on a Single RTX 3090 via vLLM — Altruistic_Heat_9531 · 2026-09-14
- Nvidia's $59.7B Quarterly Net Income Works Out to ~$656M Profit Per Day — himanshustwts · 2026-09-14
- Tiny Neural Nets Revival: Transformer Hits ~1500 tok/s on M4 CPU via SME2 Instructions — GregoryDiamos · 2026-09-14