RTX 5090 + Intel Arc B70 for local LLMs: halved throughput vs bigger context
indiealexh · reddit · 2026-10-01
A user with an RTX 5090 picked up a cheap Intel Arc B70 to extend local LLM context. Splitting a Qwen3 27B Q4KXL model across both GPUs roughly halves throughput due to the B70's lower memory bandwidth. He asks whether to run a smaller model on the B70 instead or accept slower inference for larger context, and wants advice from others with mismatched GPU setups.
More from Infra
- M5 Ultra Shootout: Qwen3.8-Flash-Next Prefills 200K in 47s vs Laguna's 455s — nonlinearsystems · 2026-10-01
- Azure AI kicks off series: why content extraction matters more as GenAI models get stronger — adnan_hashmi · 2026-10-01
- Micron Crushes Estimates as Quarterly Revenue Nearly Quadruples to $54.2B — Polymarket · 2026-10-01
- Google's compute hunger games: sells 1M TPUs to Anthropic while renting SpaceX Blackwell at 2x market — zacharynado · 2026-10-01
- Micron plans to return 100% of excess cash to shareholders from Dec 2026 — firstadopter · 2026-10-01
- Ex-OpenAI policy VP: compute is now split into ~3 tiers and the gap is widening — Miles_Brundage · 2026-10-01