Mixing a 3090 With RTX 50 Cards for Local LLM Inference: The Performance Trade-offs
mrgreatheart · reddit · 2026-09-06
A Reddit user weighing local inference hardware upgrades is stuck on mixed-architecture GPUs: they run a 5070 Ti plus two 16GB 5060 Ti cards, getting 60 tok/s decode and 1,500 pp on Qwen3.8-27B-IQ4-XS-MTP via tensor parallelism in llama.cpp.
The plan was a second 5070 Ti for two 32GB pairs, but local NVIDIA prices just jumped 25%. A 3090 Founders Edition is cheaper and offers 8GB more VRAM with slightly higher memory bandwidth — yet mixing Ampere with Blackwell reportedly hurts llama.cpp tensor-parallel performance, and vLLM effectively can't do TP with mismatched cards.
Open questions: is 8GB extra VRAM worth the performance hit, would pipeline parallelism (layer split) with two high-bandwidth cards be faster, and could the 3090 serve as the primary card holding the K/V cache?
More from Infra
- Cathie Wood claims Anthropic pays ~$50B per gigawatt while xAI's build costs sit in the high-20s — Kyrannio · 2026-09-06
- Nvidia invests $3.5B in MediaTek as chipmaker joins NVLink Fusion ecosystem — Beth_Kindig · 2026-09-06
- Duplicating one Pinokio app and deduping saved 140GB on disk — cocktailpeanut · 2026-09-06
- AWS to build 50MW AI Zone in Saudi Arabia with HUMAIN, mixing Trainium and Nvidia GPUs — Beth_Kindig · 2026-09-06
- GPT-3 to GPT-6: OpenAI API prices dropped up to 6x in five years — BLUECOW009 · 2026-09-06
- Researcher calls out Benn Jordan's "data center infrasound" health claims as nonsense — CharlesFLehman · 2026-09-06