LessThink-Qwen3-4B: post-trained to spend 44% fewer reasoning tokens on one GPU
stey1r · reddit · 2026-09-30
A developer released LessThink-Qwen3-4B, a post-trained version of Qwen3-4B that spends 44% fewer tokens on reasoning while keeping its knowledge and answer style. The entire training pipeline ran on a single GPU.
Details and the model are available on the author's project page — a useful reference for cheaply cutting local inference overhead.
More from Models
- Claim: Opus 5.5 one-shot a full music video entirely in code, no video model — repligate · 2026-09-30
- Reddit user reports receiving $62,500 in surprise ChatGPT API credits — KeyBaker5 · 2026-09-30
- dots local agent tasks: why do different models burn quota at wildly different rates? — lxfater · 2026-09-30
- User cancels ChatGPT for Claude, citing stingy limits and quota-reset tactics — ezshine · 2026-09-30
- Claude has gotten a lot better at creating infographics over the past few months — tom_doerr · 2026-09-30
- DepthBench compares 10 architectures to find which residual tweaks actually buy computational depth — SonglinYang4 · 2026-09-30