Qwen3.6-35b-a3b hits 100 tok/s on two B60 XPU cards with SYCL kernels
QuixiAI · x · 2026-07-22
- QuixiAI says it is running Qwen3.6-35b-a3b nvfp4 at 100 tokens/s.
- The setup uses 2× B60 XPU and custom SYCL kernels.
- This is a concrete performance datapoint showing that a fairly large Qwen model can be made to run efficiently on Intel-style XPU hardware.
Related event: QuixiAI Enables Cross-Backend Execution and Boosts CPU Inference(3 posts)→
More from Infra
- TokenSwitch routes tasks to cheaper models to cut token costs in Codex workflows — DeryaTR_ · 2026-07-22
- Voice agents are pushing LangChain tracing and observability beyond text workflows — Hacubu · 2026-07-22
- Lazy imports cut a Python agent’s boot time 46% and exposed a hidden import cycle — Federal-Teaching2800 · 2026-07-22
- Flyin pitches an 8-channel CCWDM module for data centers, 5G and fiber networks — glenbeer · 2026-07-22
- World Monitor turns 500+ feeds into a local AI global-intelligence dashboard — Roger_M_Taylor · 2026-07-22
- Open-source CLI proxy claims to cut Claude Code token use by 90% — Roger_M_Taylor · 2026-07-22