TensorSharp MoE Offload Benchmark: Up to 8x Faster Than llama.cpp

fuzhongkai · reddit · 2026-08-06

TensorSharp's MoE CPU-offload feature has been merged into the main branch, allowing routed expert weights to be kept in system RAM. This enables running 35B-A3B MoE models alongside long-context KV caches on 12-16GB GPUs.

The author benchmarked TensorSharp against llama.cpp across various offload depths on dual RTX PRO 6000 Blackwell GPUs:

Although TensorSharp has slightly higher VRAM usage, it achieves overwhelming inference speed advantages in MoE CPU offloading scenarios.

Related event: TensorSharp MoE Offload Slashes VRAM and Outperforms llama.cpp(2 posts)→

Original post →

More from Infra

Infra channel →