Qwen 27B Hits 1253 tok/s on 2x RTX 5090, 8-10x Cheaper Than API Pricing

me_broke · reddit · 2026-09-19

A team optimizing Qwen 3.8 27B for their inference project reports 1253 tok/s and sub-2s average TTFT on 2x RTX 5090 using sglang + Dflash2 with agent-tuned configs. At $0.88/hr on vast.ai ($650/month), it delivers inference worth $8-10k in API prices — 8-10x cheaper than OpenRouter's cheapest provider.

Original post →

More from Infra

Infra channel →