TensorSharp Adds Multi-GPU Tensor Parallelism, Boosting Local GGUF Inference Speeds
fuzhongkai · reddit · 2026-07-30
TensorSharp, an open-source native .NET inference engine, now supports Megatron-style tensor parallelism across multiple GPUs for local GGUF models.
- Core Features: Supports CUDA, Vulkan, and Metal backends, featuring continuous batching, speculative decoding, and multimodal capabilities. It works with direct CUDA, GGML CUDA/Vulkan, and multi-node setups.
- Benchmarks: On 2× RTX 2000 Ada 16GB GPUs without NVLink, prefill speeds for Gemma 4 26B-A4B jumped from 1845 to 2537 tok/s. The massive Qwen 3.5 35B-A3B, previously unfittable on a single card, achieved 18.1 tok/s decode.
- Roadmap: The developer is optimizing Qwen performance on multi-GPU setups and plans to add DeepSeek V4 support soon.
More from Infra
- Samsung Earnings Call Reveals 2nm Projects from Major CSP and HPC Customers — zephyr_z9 · 2026-07-30
- Samsung Q2 Call: Agentic AI Triggers Memory Shortage and Compute Spillover — zephyr_z9 · 2026-07-30
- NVIDIA Open-Sources PyCuTe: Pure Python Layout Algebra for CUTLASS — asdf1234_0 · 2026-07-30
- UC Berkeley's K-search: Auto-Translating CUDA Kernel Optimizations to Apple's MLX — berkeley_ai · 2026-07-30
- Advantech Edge Device Powered by Nvidia Thor Runs RealSense GMSL Cameras — chrismatthieu · 2026-07-30
- Together Offers Lowest Price and Highest Cache Hit Rate for Kimi K3 on OpenRouter — zhyncs42 · 2026-07-30