Running at 100 tok/s on Arc Pro B60

QuixiAI · x · 2026-07-19

On 2× Arc Pro B60 using QuixiCore SYCL kernel, running 35B-A3B NVFP4 achieves a speed of 100 tok/s.

This post emphasizes the author's inference optimization capability on specific hardware, not the model itself; the selling point is Arc GPU + SYCL kernel + quantized inference throughput.

Related event: Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference(2 posts)→

Original post →

More from Infra

Infra channel →