Compute Benchmarks: RTX 5090 vs 6000 PRO
panchovix · reddit · 2026-07-12
The author provides a detailed performance comparison of the RTX 5090, 6000 PRO MaxQ, and 6000 PRO WS/SE, covering image generation and local LLM inference.
They performed a shunt mod on the 6000 PRO MaxQ, combined with water cooling, pushing the power limit to around 600W, with temperatures ranging between 45°C and 60°C. A rented 6000 PRO WS from runpod was used as a baseline.
Test software and setups include:
- Various PyTorch versions
- SageAttention 2.1
- Forge neo
- RTX upscaling extensions and additional sampler extensions
- torch compile using max autotune without cudagraphs
Image generation tests were run with fixed samplers, steps, and prompts. For LLMs, llama.cpp was used with partial model offloading to the CPU, bottlenecking the main GPU to test local inference performance. Overall, the post highlights how different GPUs, power limits, and cooling solutions impact AI compute throughput.
Related event: Performance Comparison: RTX 5090 vs RTX 6000 PRO(2 posts)→
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11