Bittensor miners squeeze up to 54% more inference throughput from the same GPUs

bittingthembits · x · 2026-10-11

Bittensor SN10 (Pareton) miners benchmarked Qwen3.8-27B with SGLang on 4x RTX 5090s using NVIDIA's AIPerf, completing 600/600 requests with same-model, same-GPU throughput gains of +54% at 4 concurrent requests, +50% at 8, +40% at 16, +23% at 32 — roughly 35% lower compute cost per token by the poster's math.

How it works: miners compete to optimize the model-serving software; SN10 tests each submission on the same workload and hardware, keeping only optimizations that are faster without hurting output quality. Each winner becomes the next baseline — a ratchet that only goes up. Per opentensor, a miner-developed optimization has already been merged into vLLM.

Original post →

More from Infra

Infra channel →