New RL Trick: Teaching Models to Save Reasoning Tokens
andrey_kurenkov · x · 2026-07-17
Thinking Machines Inkling proposed a practical Reinforcement Learning (RL) trick: incorporating a penalty term into the optimization objective, specifically `Reward = Task Reward − λ × (number of reasoning tokens)`, alongside task success rate. By adjusting the λ value across different rollouts and pairing it with varying effort instructions, the model learns to treat "reasoning" as a resource to be conserved, rather than endlessly maximizing its thinking length.
More from Research
- Knowledgeless Language Models cut closed-book recall by anonymizing entities during pretraining — gdm3000 · 2026-07-21
- CPU-native LLM pilot passes 4 of 5 gates, but cross-tokenizer distillation still loses — WildPino25 · 2026-07-21
- A GPT 5.6 Sol workflow reportedly generates an infinite family of counterexamples — OwariDa · 2026-07-21
- A research guide v7 surfaces two contradictions instead of smoothing them over — Fantastic_Aside6599 · 2026-07-21
- Agents can remember facts, but still forget how to do the job — No_Advertising2536 · 2026-07-21
- AI-assisted search finds small counterexamples to the Gaussian Moments Conjecture — RichmanRonald · 2026-07-21