RTX 5090 Inference Test: Capping Power at 480W Costs Less Than 3% Performance
WonderfulEagle7096 · reddit · 2026-08-04
Addressing heat and noise concerns with the RTX 5090 during local inference, a developer's tests reveal that limiting the power draw to 480W is the optimal sweet spot.
Benchmarking with Qwen 3.6-27B (Q6K), capping at 480W results in only a 2.1% drop in decode tokens/sec and an 8.8% drop in prefill speed compared to the default 566W. This significantly reduces fan noise and heat while increasing card longevity. Pushing it down to 450W causes performance to fall off much faster (4.2% decode loss).
More from Infra
- Open-Source 'Personal AI Computer': Build a Local AI Data Center with RTX 5090s — dee_hw · 2026-08-05
- AI Compute Demand to Quadruple Data Center Electricity by 2030 — bittingthembits · 2026-08-05
- DeepSeek V4 Flash Local Deployment Hits 16k Output Limit — El_90 · 2026-08-05
- China's Potential InP Wafer Export Ban Threatens Global AI Infrastructure — teortaxesTex · 2026-08-05
- 3-Hour Deep Dive with vLLM Core Dev: Open Source Infra and Model Co-design — vista8 · 2026-08-05
- Running MiniMax H3 on a Single 3090: Troubleshooting OOMs and Optimizing Args — knoll_gallagher · 2026-08-05