On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding?
MasterNomie · reddit · 2026-09-28
A local-LLM user on RTX 5090 + 96GB DDR5 reports Qwen 3.8 27B decoding at 200+ TPS versus 50 TPS for Flash next, and asks for a coding model in between that sustains 75-100 TPS. Failing that, they propose splitting the workflow: Flash next for planning, 27B for implementation, with a RAM upgrade to 128GB planned.
More from Infra
- Lightning AI and Google Cloud cut PyTorch Lightning checkpoint write times by up to 95% — LightningAI · 2026-09-28
- Crucible Capital founder financed her own GPU cluster with stablecoin debt and personal credit risk — MarvinTBaumann · 2026-09-28
- NVIDIA launches Open Agent Safety Platform for controlling what AI agents can do — nvidia · 2026-09-28
- MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul — awnihannun · 2026-09-28
- Cloudflare incident: skipped block zeroing leaked tenant data across 18 of 24 containers — arpit_bhayani · 2026-09-28
- Lumen Launches On-Demand Dedicated Internet Up to 100 Gbps at 10M US Sites — shashib · 2026-09-28