LessThink-Qwen3-4B: post-trained to spend 44% fewer reasoning tokens on one GPU

stey1r · reddit · 2026-09-30

A developer released LessThink-Qwen3-4B, a post-trained version of Qwen3-4B that spends 44% fewer tokens on reasoning while keeping its knowledge and answer style. The entire training pipeline ran on a single GPU.

Details and the model are available on the author's project page — a useful reference for cheaply cutting local inference overhead.

Original post →

More from Models

Models channel →