TensorSharp Beats llama.cpp in DeepSeek V4 Flash Inference Benchmark

fuzhongkai · reddit · 2026-08-01

Open-source inference engine TensorSharp has added support for the DeepSeek-V4-Flash-0731 model, outperforming llama.cpp in multi-GPU benchmarks.

Tested on 4x Nvidia A40 GPUs (CUDA 12.8) using unsloth's Q8KXL quantized version:

The engine supports CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, and multi-GPU/node inference.

Related event: TensorSharp Outperforms llama.cpp in Multi-GPU DeepSeek Inference(2 posts)→

Original post →

More from Infra

Infra channel →