H200 vs multi-GPU RTX PRO 6000 Blackwell: how to pick inference hardware by budget

recentheartbroken · reddit · 2026-10-07

A hands-on comparison for building an inference server on equal budgets. RTX PRO 6000: 96GB GDDR7 at 1.792 TB/s, native FP4, no NVLink; H200: 141GB HBM3e at 4.8 TB/s with NVLink, FP8 minimum.

Key caveat: decode is usually memory-bandwidth-bound, not compute-bound, so datasheet TFLOPS says almost nothing about tokens/sec.

Original post →

More from Infra

Infra channel →