Microsoft's When2Think boosts AIME24 Pass@3 by 10% while cutting tokens 27.9%
microsoft · hf · 2026-09-18
Microsoft proposes When2Think, a post-training framework for hybrid reasoning that tackles the systematic inefficiency of large reasoning models: overthinking easy problems and underthinking hard ones. It formulates efficient reasoning as instance-adaptive compute allocation, using Instance-level Difficulty-Aware Control (IDAC) reward shaping with pre-computed accuracy/token reference statistics, enabling stable critic-free optimization without learned reward models. On AIME24, Pass@3 rises 10.0% while token usage drops 27.9%; on AIME25 it reaches 40.0% Pass@3, beating compression- and routing-only baselines.
More from Research
- Turning Yang-Mills existence and mass gap into a formal conjecture is AI math's ultimate test — geoffreyirving · 2026-09-18
- CoRL 2026 SPIN workshop pits robots against a child in live dexterity challenge — berkeley_ai · 2026-09-18
- That viral "air-gapped data exfiltration" paper only read temperature over a 4cm gap — basedjensen · 2026-09-18
- An AI forecaster has won the seasonal Metaculus Cup for the first time — NathanpmYoung · 2026-09-18
- New video: weather forecasting with neural networks explained — ariG23498 · 2026-09-18
- First LLM runs entirely on-chain: weights, activations, attention all in SVM transactions — AccBalanced · 2026-09-18