A Qwen 3.6 27B derivative claims to cut thinking tokens by more than 90%
AppealSame4367 · reddit · 2026-07-23
A Qwen 27B derivative claims 90% fewer thinking tokens
A Reddit user spotted ProCreations/grug-27b on Hugging Face and says the benchmark claims are notable:
- It is presented as a Qwen 3.6 27B derivative
- The repo claims it can use more than 90% fewer tokens for the thinking part
- If true, the poster says it would make a 27B model running at 3 TPS on an old laptop feel much faster in practice
The user has not tested it yet, so the post is still a claim-plus-curiosity rather than a verified evaluation.
More from Models
- Claude users warned usage limits may reset if Opus 5 launches today — CtrlAltDwayne · 2026-07-23
- X rumor says GPT-5.6, Cerebras 750 token/s release and Claude Opus 5 may land today — Scobleizer · 2026-07-23
- OpenAI and Codex climb a Chinese trend tracker as attention rises — huangyun_122 · 2026-07-23
- Google's Gemini 3.5 Flash Positioned as a Cost-Effective Workhorse Model — koltregaskes · 2026-07-23
- Another reply says Grok still failed to count the teams correctly — ivan_bezdomny · 2026-07-23
- A user says Grok 4.5 High is now their daily go-to over Claude — prasenx · 2026-07-23