Microsoft's When2Think boosts AIME24 Pass@3 by 10% while cutting tokens 27.9%

microsoft · hf · 2026-09-18

Microsoft proposes When2Think, a post-training framework for hybrid reasoning that tackles the systematic inefficiency of large reasoning models: overthinking easy problems and underthinking hard ones. It formulates efficient reasoning as instance-adaptive compute allocation, using Instance-level Difficulty-Aware Control (IDAC) reward shaping with pre-computed accuracy/token reference statistics, enabling stable critic-free optimization without learned reward models. On AIME24, Pass@3 rises 10.0% while token usage drops 27.9%; on AIME25 it reaches 40.0% Pass@3, beating compression- and routing-only baselines.

Original post →

More from Research

Research channel →