Qwen 27B NVFP4 on RTX 5090: 120 t/s with vision and 451K cache
t4a8945 · reddit · 2026-08-23
A detailed technical report on running Qwen 2.5 27B in NVFP4 quantization on a single RTX 5090 (400W limited). Using vLLM, the setup achieves 120 t/s with vision enabled and 451K global KV-cache. Includes extensive benchmarks from 4K to 185K context and a setup guide.
More from Infra
- Singapore launches server powered by living human brain cells with 100x energy efficiency — evilsocket · 2026-08-23
- Best Practice for Building AI Production Systems: From Inference to Deployment — _ScottCondron · 2026-08-23
- Codex Rate Limits Breaking Existing Workflows, Real Token Costs Incoming — StewartalsopIII · 2026-08-23
- ShardFlow: 28 TPS on Qwen2.5-7B over WAN via speculative decoding — katua_bkl · 2026-08-23
- Cursor Launches S3-Based Git Storage System for Scale — bibryam · 2026-08-23
- PromoteOps: MCP Server for Automating AWS CloudFormation Promotion — iamthanoss · 2026-08-23