When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
The paper addresses the inefficiency of Large Reasoning Models, which often overthink easy problems and underthink hard ones. The authors propose When2Think, a post-training framework that dynamically allocates computation based on problem difficulty. The method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism leveraging pre-computed reference statistics. Experimental results on AIME24 show that Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model.
On AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model.
On AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.
IDAC enables stable critic-free optimization without learned reward models or online reference-model queries.