Qwen3.5 35B A3B hits 55 tok/s on an RTX 5060 Ti with an extended Garlic build

Azazelionide · reddit · 2026-07-24

A user reports running Qwen3.5 35B A3B float8 at 55 tokens/s on an RTX 5060 Ti using an extended version of Garlic.

They say the result comes from some Gated Delta Network kernel work, and that it significantly outperforms llama.cpp running the same model in Q8 quantization. The reported number drops to 61 tok/s without recording because screen recording consumes CPU/GPU resources.

They also note that this is without MTP, and that MTP could speed generation up further. The author plans to write a blog post explaining the trick behind the speedup.

Original post →

More from Infra

Infra channel →