Running at 100 tok/s on Arc Pro B60
QuixiAI · x · 2026-07-19
On 2× Arc Pro B60 using QuixiCore SYCL kernel, running 35B-A3B NVFP4 achieves a speed of 100 tok/s.
This post emphasizes the author's inference optimization capability on specific hardware, not the model itself; the selling point is Arc GPU + SYCL kernel + quantized inference throughput.
Related event: Dual Arc Pro B60 Achieves 100 tok/s on 35B Model Inference(2 posts)→
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21