Qwen3.5 35B A3B hits 55 tok/s on an RTX 5060 Ti with an extended Garlic build
Azazelionide · reddit · 2026-07-24
A user reports running Qwen3.5 35B A3B float8 at 55 tokens/s on an RTX 5060 Ti using an extended version of Garlic.
They say the result comes from some Gated Delta Network kernel work, and that it significantly outperforms llama.cpp running the same model in Q8 quantization. The reported number drops to 61 tok/s without recording because screen recording consumes CPU/GPU resources.
They also note that this is without MTP, and that MTP could speed generation up further. The author plans to write a blog post explaining the trick behind the speedup.
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27