Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference
A deployment demonstration shows that running a 35B-A3B model in NVFP4 precision on two Intel Arc Pro B60 GPUs with the QuixiCore SYCL kernel achieves an inference speed of 100 tokens per second.
2026-07-19 ~ 2026-07-19 · 2 related posts
- Running at 100 tok/s on Arc Pro B60 — QuixiAI · 2026-07-19
- 256k Context Inference Hits 100 tok/s — QuixiAI · 2026-07-19