Reddit user runs Nemotron Ultra 550B across aging MI50 and P40 GPU rigs
Old_Grapefruit8774 · reddit · 2026-07-22
A Reddit user reports running Nemotron Ultra 550B IQ3S across two aging GPU boxes linked by 100GbE and shares throughput results, hardware specs, and the exact launch command.
- One machine uses 7× MI50s + 2× MI50s with a total of 176 GB VRAM; the other uses 5× P40s with 120 GB VRAM.
- The post includes benchmark tables for pp512, tg128, and combined context lengths up to 126,720 tokens.
- The author says the results are surprisingly good on old hardware and is now considering adding Chinese 22GB RTX 2080 Ti cards to expand VRAM further.
More from Infra
- A hybrid local-plus-cloud inference model is the AI equivalent of 65 MPH driving — dmitry140 · 2026-07-22
- Alphabet capex call may matter less than what the spending is buying — tengyanAI · 2026-07-22
- U.S. firms may need Chinese models for cyber defense, yet one breach could trigger a ban — natolambert · 2026-07-22
- DeepSeek’s hardware quadrant would be notable, and Alibaba also has its own chips — teortaxesTex · 2026-07-22
- Triton’s TTGIR layer drives both most of its value and many performance bugs — dhruv2038 · 2026-07-22
- Do prompt caches meaningfully cut costs for production AI agents? — MembershipEmergency7 · 2026-07-22