Qwen3.5 35B A3B hits 55 tok/s on an RTX 5060 Ti with an extended Garlic build
Azazelionide · reddit · 2026-07-24
A user reports running Qwen3.5 35B A3B float8 at 55 tokens/s on an RTX 5060 Ti using an extended version of Garlic.
They say the result comes from some Gated Delta Network kernel work, and that it significantly outperforms llama.cpp running the same model in Q8 quantization. The reported number drops to 61 tok/s without recording because screen recording consumes CPU/GPU resources.
They also note that this is without MTP, and that MTP could speed generation up further. The author plans to write a blog post explaining the trick behind the speedup.
More from Infra
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11