TensorSharp MoE Offload Benchmark: Up to 8x Faster Than llama.cpp
fuzhongkai · reddit · 2026-08-06
TensorSharp's MoE CPU-offload feature has been merged into the main branch, allowing routed expert weights to be kept in system RAM. This enables running 35B-A3B MoE models alongside long-context KV caches on 12-16GB GPUs.
The author benchmarked TensorSharp against llama.cpp across various offload depths on dual RTX PRO 6000 Blackwell GPUs:
- Gemma 4 26B: With full offloading, TensorSharp significantly leads in both prompt processing (pp) and generation (tg) speeds, reaching up to 5.59x and 3.07x faster respectively.
- Qwen 3.5 35B: The speed advantage is even more pronounced, with pp speed up to 8.85x faster and tg speed up to 4.50x faster on full offload.
- GPT-OSS 20B: TensorSharp maintains a 5-7x advantage in pp speed and nearly 3x in tg speed during offloaded scenarios.
Although TensorSharp has slightly higher VRAM usage, it achieves overwhelming inference speed advantages in MoE CPU offloading scenarios.
Related event: TensorSharp MoE Offload Slashes VRAM and Outperforms llama.cpp(2 posts)→
More from Infra
- Developers Mock Google's Squandered GPU Ecosystem Advantage and TPU Hype — teortaxesTex · 2026-08-06
- DeepInfra Serves 700B+ Tokens/Day on NVIDIA Blackwell Ultra B300s — gharik · 2026-08-06
- Hippius on Bittensor: Decentralized S3 Storage with ~900TB Capacity — markjeffrey · 2026-08-06
- Qdrant 1.19 Released: Introduces TurboQuant & Memory Tiers — qdrant_engine · 2026-08-06
- Marvell Photonic Fabric Wins AI Infrastructure Award for Scaling Inference — BenBajarin · 2026-08-06
- Aeva Pivots to AI Data Center Optical Interconnects with Hyperscaler Deal — BenBajarin · 2026-08-06