Mixing a 3090 with an Intel Arc B70 for 56GB VRAM local LLM inference?
overand · reddit · 2026-09-15
Running local LLMs on a 3090 (24GB) + 4070 Ti (12GB), the poster considers adding an Intel Arc Pro B70 to reach 56GB VRAM after one 3090 started failing. They know llama.cpp can tensor-split across cards via Vulkan or dual backends (CUDA + OpenVINO/SYCL with llama-rpc), but worry about configuration pain and performance loss. Full host specs included: Ryzen 5 3600, 128GB DDR4 (one pair with bad blocks stabilized via lower clocks + badram kernel line), Ubuntu Server 24.04. The thread centers on the feasibility and pitfalls of heterogeneous-GPU inference.
More from Infra
- OpenAI engineers: kernel optimization cut GPT-5.6 Sol serving cost by 20% — TheTuringPost · 2026-09-15
- Devin left alone with Modal H100s cuts training kernel peak memory 46% and latency 52% — AAAzzam · 2026-09-15
- Why Macs quietly win at local AI: unified memory beats RTX 5090 and accessibility APIs power better computer use — dotey · 2026-09-15
- 5 production apps, 2M monthly requests for $6: why developers are going all-in on Cloudflare — viksit · 2026-09-15
- Google Cloud and Inferact partner to make TPU a first-class vLLM target — vllm_project · 2026-09-15
- How Abnormal AI Screens Billions of Emails with Bedrock AgentCore Sandboxes — AWS ML Blog · 2026-09-15