Mixing a 3090 With RTX 50 Cards for Local LLM Inference: The Performance Trade-offs

mrgreatheart · reddit · 2026-09-06

A Reddit user weighing local inference hardware upgrades is stuck on mixed-architecture GPUs: they run a 5070 Ti plus two 16GB 5060 Ti cards, getting 60 tok/s decode and 1,500 pp on Qwen3.8-27B-IQ4-XS-MTP via tensor parallelism in llama.cpp.

The plan was a second 5070 Ti for two 32GB pairs, but local NVIDIA prices just jumped 25%. A 3090 Founders Edition is cheaper and offers 8GB more VRAM with slightly higher memory bandwidth — yet mixing Ampere with Blackwell reportedly hurts llama.cpp tensor-parallel performance, and vLLM effectively can't do TP with mismatched cards.

Open questions: is 8GB extra VRAM worth the performance hit, would pipeline parallelism (layer split) with two high-bandwidth cards be faster, and could the 3090 serve as the primary card holding the K/V cache?

Original post →

More from Infra

Infra channel →