Mixing a 3090 with an Intel Arc B70 for 56GB VRAM local LLM inference?

overand · reddit · 2026-09-15

Running local LLMs on a 3090 (24GB) + 4070 Ti (12GB), the poster considers adding an Intel Arc Pro B70 to reach 56GB VRAM after one 3090 started failing. They know llama.cpp can tensor-split across cards via Vulkan or dual backends (CUDA + OpenVINO/SYCL with llama-rpc), but worry about configuration pain and performance loss. Full host specs included: Ryzen 5 3600, 128GB DDR4 (one pair with bad blocks stabilized via lower clocks + badram kernel line), Ubuntu Server 24.04. The thread centers on the feasibility and pitfalls of heterogeneous-GPU inference.

Original post →

More from Infra

Infra channel →