ThinkingCap-Qwen3.6-27B delivers 35–45 tok/s with no obvious quality loss
TinyFrodo · reddit · 2026-07-28
A user says ThinkingCap-Qwen3.6-27B F16 has become their default replacement for Qwen3.5-27B F16 after two days of use. They report roughly 35–45 tok/s versus 30–40 tok/s before, with no noticeable quality drop, and suspect reduced token usage is helping. The post shares a full llama.cpp server command, including spec-draft-n-max 4, and asks for further ways to push throughput higher.
More from Infra
- Kimi K3 is an API-only multi-node model for now; local parity may take months to years — johnseach · 2026-07-28
- Only one of 63 API companies exposes all three agent-ready surfaces — ExtensionPea834 · 2026-07-28
- Reddit asks whether 136.7 million x402 settlements actually prove agent adoption — heiba_wk · 2026-07-28
- Meme argues local models could cut middlemen and lower hardware demand — HugoCortell · 2026-07-28
- Builder says a managed graph DB would cost $25K for 16GB, so he made one for Apple Silicon — arthurcolle · 2026-07-28
- Moonshot’s Kimi K3 lands on Nebius as an OpenAI-compatible API with 1M context — rohanpaul_ai · 2026-07-28