RTX 5090 Inference Test: Capping Power at 480W Costs Less Than 3% Performance

WonderfulEagle7096 · reddit · 2026-08-04

Addressing heat and noise concerns with the RTX 5090 during local inference, a developer's tests reveal that limiting the power draw to 480W is the optimal sweet spot.

Benchmarking with Qwen 3.6-27B (Q6K), capping at 480W results in only a 2.1% drop in decode tokens/sec and an 8.8% drop in prefill speed compared to the default 566W. This significantly reduces fan noise and heat while increasing card longevity. Pushing it down to 450W causes performance to fall off much faster (4.2% decode loss).

Original post →

More from Infra

Infra channel →