256k Context Inference Hits 100 tok/s

QuixiAI · x · 2026-07-19

Showcases an inference deployment result: running **NVFP4** inference for the **35B-A3B** model at **100 tok/s** using the **QuixiCore SYCL kernel** on **2x Arc Pro B60** GPUs. The focus here isn't on the model itself, but rather that it provides a concrete use case combining **long context (256k) + VRAM/hardware + kernel optimization**, serving as a solid reference for local/edge inference and performance tuning.

Related event: Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference(2 posts)→

Original post →

More from Infra

Infra channel →