Intel Arc Pro B70 Benchmarks: 52 tok/s with Qwen3.8-27B via vLLM XPU
No-Dot-Not · reddit · 2026-08-22
The author benchmarks the Intel Arc Pro B70 (32GB) running Qwen3.8-27B INT4 on Linux. Using vLLM XPU with MTP2 speculative decoding, it achieves 52.2 tok/s median decode speed at 64K context, outperforming llama.cpp SYCL by 1.8x. The test covers 128K context validation, vision capabilities, tool calling, and a real-world coding agent task. The post details current XPU stack bugs (e.g., concurrency limits with MTP) and concludes the B70 is a viable 32GB inference card with the right configuration.
More from Infra
- Proposal to convert offshore oil rigs into data centers — beffjezos · 2026-08-22
- OpenAI DNS Records Hint at Parallel Banking and Hardware Infrastructure — imjustnewatai · 2026-08-22
- AI energy crisis: ChatGPT queries use 10x energy of Google search — ingliguori · 2026-08-22
- Google Cloud launches Global Front End for cross-cloud networking — rseroter · 2026-08-22
- Opinion: 'Ban Data Centers' is a Luxury Belief That Would Disastrously Impact Economy — robleclerc · 2026-08-22
- Path to 100x AI Efficiency? Needs Architecture Shift, Says VC — prateekj · 2026-08-22