Intel B70 users report broken multi-GPU llama.cpp scaling and garbled output
nick_ziv · reddit · 2026-07-24
A user reports trying to run llama.cpp on Intel B70 GPUs under Ubuntu 26.04 and finding that multi-GPU scaling behaves strangely.
Their numbers show:
- CYSL with 1 GPU: 650 prefill, 24 decode
- CYSL with 2 GPUs: 750 prefill, but garbled decode at 23 t/s
- Vulkan with 1 GPU: 450 prefill, 20 t/s decode
- Vulkan with 2 GPUs: 350 prefill, 18 t/s decode
- A 6-GPU Qwen 3.5 122B run on A10B/Vulkan: 160 prefill and 9 t/s decode
They suspect something is wrong with the stack and say they now regret buying the B70s, asking for help or alternative suggestions.
More from Infra
- Sam Altman says 750 token/sec is coming to 5.6 sol in July — ChrisGPT · 2026-07-24
- Turso bets on BYOC databases running entirely on S3 in sovereign clouds — glcst · 2026-07-24
- Agent Substrate targets sandboxed agents with Kubernetes-style orchestration — bibryam · 2026-07-24
- A new 4,000-GPU RTX 6000 Pro Blackwell cluster is coming online for burn-in — isidentical · 2026-07-24
- Intel’s order book could make 2027 capex much higher than investors expect — BenBajarin · 2026-07-24
- Intel signals a move back into part of the memory business — BenBajarin · 2026-07-24