RTX 5090 + Intel Arc B70 for local LLMs: halved throughput vs bigger context

indiealexh · reddit · 2026-10-01

A user with an RTX 5090 picked up a cheap Intel Arc B70 to extend local LLM context. Splitting a Qwen3 27B Q4KXL model across both GPUs roughly halves throughput due to the B70's lower memory bandwidth. He asks whether to run a smaller model on the B70 instead or accept slower inference for larger context, and wants advice from others with mismatched GPU setups.

Original post →

More from Infra

Infra channel →