A Qwen 3.6 27B derivative claims to cut thinking tokens by more than 90%
AppealSame4367 · reddit · 2026-07-23
A Qwen 27B derivative claims 90% fewer thinking tokens
A Reddit user spotted ProCreations/grug-27b on Hugging Face and says the benchmark claims are notable:
- It is presented as a Qwen 3.6 27B derivative
- The repo claims it can use more than 90% fewer tokens for the thinking part
- If true, the poster says it would make a 27B model running at 3 TPS on an old laptop feel much faster in practice
The user has not tested it yet, so the post is still a claim-plus-curiosity rather than a verified evaluation.
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11