Qwen3.6 Hits 65 tok/s on Dual B60s

QuixiAI · x · 2026-07-10

A user ran nvidia/Qwen3.6-35B-A3B-NVFP4 on dual B60s, optimized with a custom SYCL kernel. Tests show a context length of up to 128k and an inference speed of 65 tok/s, with DFlash disabled.

Original post →

More from Infra

Infra channel →