Mixing 3x RTX 5090 with AMD GPUs for DeepSeek: A Local Rig Experiment
fluffywuffie90210 · reddit · 2026-08-10
An AI hardware enthusiast shares their experience expanding VRAM by integrating an AMD GPU into a 3x RTX 5090 system.
- Hardware Setup: Plans to use a PCIe riser cable to add an AMD 9700 AI Pro, aiming for a 128GB VRAM dream machine specifically to run massive DeepSeek models locally.
- Technical Bottlenecks: Highlights compatibility issues mixing CUDA and ROCm. While workarounds exist using an RPC server with llama.cpp, MoE (Mixture of Experts) models suffer terrible acceleration over multi-card RPC setups—only dense models see good speedups.
- Cost Trade-offs: The author is torn between spending heavily to maximize VRAM or settling for a cheaper 5070Ti for gaming, turning to the community for advice.
More from Infra
- GPU Hot: Lightweight Self-Hosted Real-Time NVIDIA GPU Dashboard — tom_doerr · 2026-08-10
- Google Open-Sources TPU Raiden Inference Library for KVCache Transfer — xennygrimmato_ · 2026-08-10
- AI compute becomes strategic as tech giants pledge to build their own power infrastructure — bittingthembits · 2026-08-10
- DwarfStar Accelerates DeepSeek Inference with DFlash Speculative Decoding — antirez · 2026-08-10
- Neural AI Breakthrough: Memory Chip Reconstructs Human Cortex in Real Time — Dr_Alex_Crimi · 2026-08-10
- Cursor and Together AI Partner for Low-Latency AI Coding Inference — togethercompute · 2026-08-10