Intel Arc Pro B70 Benchmarks: 52 tok/s with Qwen3.8-27B via vLLM XPU

No-Dot-Not · reddit · 2026-08-22

The author benchmarks the Intel Arc Pro B70 (32GB) running Qwen3.8-27B INT4 on Linux. Using vLLM XPU with MTP2 speculative decoding, it achieves 52.2 tok/s median decode speed at 64K context, outperforming llama.cpp SYCL by 1.8x. The test covers 128K context validation, vision capabilities, tool calling, and a real-world coding agent task. The post details current XPU stack bugs (e.g., concurrency limits with MTP) and concludes the B70 is a viable 32GB inference card with the right configuration.

Original post →

More from Infra

Infra channel →