2x5090 over 100G RPC runs Qwen at 90-100 tok/s decode, 3000 tok/s prefill
ilarp · reddit · 2026-09-16
A user linked two single-RTX-5090 machines via ConnectX-5 100G and ran Qwen (unsloth quants, Q2) over llama RPC: 90-100 tok/s decode and 3000 tok/s prefill.
They ask for comparisons from dual-5090 or 6000 Pro owners, and admit that in practice the larger distributed setup hasn't beaten simply running the 27B model — useful data for anyone weighing distributed inference economics.
More from Infra
- Lithos: stop treating AI benchmarks as proof, define your own metrics — JiaZhihao · 2026-09-16
- Cheap third-party open LLM providers ranked, coral bricks tops the list — opensourcecolumbus · 2026-09-16
- Could an Xbox Series X cluster with 160GB unified memory host local LLMs? A Reddit thought experiment — RecursiveCTE · 2026-09-16
- Rob Mulla marks one year at Google scaling TPU inference with vLLM — Rob_Mulla · 2026-09-16
- Jev claims new frontier model 40-400x cheaper and 20-200x faster — albelfio · 2026-09-16
- Nvidia B200 Compute Prices Jump 21% in a Month as AI Demand Outstrips Supply — JOBhakdi · 2026-09-16