Bittensor miners squeeze up to 54% more inference throughput from the same GPUs
bittingthembits · x · 2026-10-11
Bittensor SN10 (Pareton) miners benchmarked Qwen3.8-27B with SGLang on 4x RTX 5090s using NVIDIA's AIPerf, completing 600/600 requests with same-model, same-GPU throughput gains of +54% at 4 concurrent requests, +50% at 8, +40% at 16, +23% at 32 — roughly 35% lower compute cost per token by the poster's math.
How it works: miners compete to optimize the model-serving software; SN10 tests each submission on the same workload and hardware, keeping only optimizations that are faster without hurting output quality. Each winner becomes the next baseline — a ratchet that only goes up. Per opentensor, a miner-developed optimization has already been merged into vLLM.
More from Infra
- DeepSeek v4.1 reportedly boosts long-context prefill throughput by ~40% — HankYeomans · 2026-10-11
- GPU Tsunami: how advanced packaging is reshaping the semiconductor test market — BenBajarin · 2026-10-11
- Optical testing is the key bottleneck for scaling co-packaged optics deployment — BenBajarin · 2026-10-11
- Cboe's former HQ to be converted into a 33MW data center with above-market rents — deanwball · 2026-10-11
- Microsoft open-sources bitnet.cpp, running 100B models on CPU at 6.17x speed — HildeKuehne · 2026-10-11
- Qualcomm CEO: AI firms want phones running 100B-parameter models continuously by 2028 — zephyr_z9 · 2026-10-11