TensorSharp Beats llama.cpp in DeepSeek V4 Flash Multi-GPU Prefill Benchmark
fuzhongkai · reddit · 2026-08-01
The open-source inference engine TensorSharp recently added support for multi-GPU and multi-node inference. The author tested its performance running the DeepSeek-V4-Flash-0731 quantized model against llama.cpp.
Test Environment & Model:
- Hardware: 4x Nvidia A40 GPUs (CUDA 12.8)
- Model: DeepSeek-V4-Flash-0731-UD-Q8KXL (from unsloth)
Benchmark Results (TensorSharp CUDA backend vs llama.cpp):
- Prefill (16K): 836 tok/s (TensorSharp optimal) vs 558 tok/s (llama.cpp)
- Decode (short): 31.5 tok/s vs 35.3 tok/s
- Decode (16K): 28.5 tok/s vs 32.2 tok/s
TensorSharp shows a significant speed advantage in long-context prefill, though it is slightly slower than llama.cpp during decoding.
Related event: TensorSharp Outperforms llama.cpp in Multi-GPU DeepSeek Inference(2 posts)→
More from Infra
- a16z: AI Infrastructure Demand Shows No Signs of Slowing Amid Supply Chain Snags — a16z · 2026-08-01
- Together AI Deep Dive: Autoscaling Endpoints for LLM Inference — togethercompute · 2026-08-01
- Vercel AI Gateway Adds Team and Project Spend Budgets — cramforce · 2026-08-01
- Tesla Signs 469MW Solar Deals to Lock in AI Compute Power Years Ahead — XFreeze · 2026-08-01
- Local Deployment on DGX Spark: Exploring Upgrades Beyond Qwen 3.5 122B — Voxandr · 2026-08-01
- Analyst Spots Equinix Expanding San Jose Campus by ~200MW — BenBajarin · 2026-08-01