ThinkingCap-Qwen3.6-27B delivers 35–45 tok/s with no obvious quality loss

TinyFrodo · reddit · 2026-07-28

A user says ThinkingCap-Qwen3.6-27B F16 has become their default replacement for Qwen3.5-27B F16 after two days of use. They report roughly 35–45 tok/s versus 30–40 tok/s before, with no noticeable quality drop, and suspect reduced token usage is helping. The post shares a full llama.cpp server command, including spec-draft-n-max 4, and asks for further ways to push throughput higher.

Original post →

More from Infra

Infra channel →