On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding?

MasterNomie · reddit · 2026-09-28

A local-LLM user on RTX 5090 + 96GB DDR5 reports Qwen 3.8 27B decoding at 200+ TPS versus 50 TPS for Flash next, and asks for a coding model in between that sustains 75-100 TPS. Failing that, they propose splitting the workflow: Flash next for planning, 27B for implementation, with a RAM upgrade to 128GB planned.

Original post →

More from Infra

Infra channel →