2x5090 over 100G RPC runs Qwen at 90-100 tok/s decode, 3000 tok/s prefill

ilarp · reddit · 2026-09-16

A user linked two single-RTX-5090 machines via ConnectX-5 100G and ran Qwen (unsloth quants, Q2) over llama RPC: 90-100 tok/s decode and 3000 tok/s prefill.

They ask for comparisons from dual-5090 or 6000 Pro owners, and admit that in practice the larger distributed setup hasn't beaten simply running the 27B model — useful data for anyone weighing distributed inference economics.

Original post →

More from Infra

Infra channel →