Qwen 27B Hits 1253 tok/s on 2x RTX 5090, 8-10x Cheaper Than API Pricing
me_broke · reddit · 2026-09-19
A team optimizing Qwen 3.8 27B for their inference project reports 1253 tok/s and sub-2s average TTFT on 2x RTX 5090 using sglang + Dflash2 with agent-tuned configs. At $0.88/hr on vast.ai ($650/month), it delivers inference worth $8-10k in API prices — 8-10x cheaper than OpenRouter's cheapest provider.
More from Infra
- AWS treats DRAM as the scarce resource in agent infra with new AgentCore runtime — bookwormengr · 2026-09-19
- Jeff Dean: RL-accelerated chip design could cut cycles from 150 people and 2 years to 10 in 3 months — haider1 · 2026-09-19
- Microsoft reportedly plans to triple data center capacity to 38GW by 2032 — Beth_Kindig · 2026-09-19
- How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — maxall4 · 2026-09-19
- India hosts ~20% of global chip design engineers, and the real share is likely higher — bookwormengr · 2026-09-19
- Micron and the AI memory cycle: HBM heads to $100B by 2027, 5-year deals reshape the trade — Beth_Kindig · 2026-09-19