TensorSharp Beats llama.cpp in DeepSeek V4 Flash Inference Benchmark
fuzhongkai · reddit · 2026-08-01
Open-source inference engine TensorSharp has added support for the DeepSeek-V4-Flash-0731 model, outperforming llama.cpp in multi-GPU benchmarks.
Tested on 4x Nvidia A40 GPUs (CUDA 12.8) using unsloth's Q8KXL quantized version:
- Prefill speed: TensorSharp's CUDA backend hit 836 tok/s, beating llama.cpp's 558 tok/s.
- Decode speed: TensorSharp also slightly outperformed llama.cpp on both short and 16K context sequences.
The engine supports CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, and multi-GPU/node inference.
Related event: TensorSharp Outperforms llama.cpp in Multi-GPU DeepSeek Inference(2 posts)→
More from Infra
- a16z: AI Infrastructure Demand Shows No Signs of Slowing Amid Supply Chain Snags — a16z · 2026-08-01
- Together AI Deep Dive: Autoscaling Endpoints for LLM Inference — togethercompute · 2026-08-01
- Vercel AI Gateway Adds Team and Project Spend Budgets — cramforce · 2026-08-01
- Tesla Signs 469MW Solar Deals to Lock in AI Compute Power Years Ahead — XFreeze · 2026-08-01
- Local Deployment on DGX Spark: Exploring Upgrades Beyond Qwen 3.5 122B — Voxandr · 2026-08-01
- Analyst Spots Equinix Expanding San Jose Campus by ~200MW — BenBajarin · 2026-08-01