Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference

A deployment demonstration shows that running a 35B-A3B model in NVFP4 precision on two Intel Arc Pro B60 GPUs with the QuixiCore SYCL kernel achieves an inference speed of 100 tokens per second.

2026-07-19 ~ 2026-07-19 · 2 related posts