CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 47 upvotes

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

QUESTION — How can the computational efficiency of Large Reasoning Models be optimized by dynamically controlling reasoning depth based on problem difficulty?

The paper addresses the inefficiency of Large Reasoning Models, which often overthink easy problems and underthink hard ones. The authors propose When2Think, a post-training framework that dynamically allocates computation based on problem difficulty. The method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism leveraging pre-computed reference statistics. Experimental results on AIME24 show that Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model.

On AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model.

On AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.

IDAC enables stable critic-free optimization without learned reward models or online reference-model queries.

junshim · 17 Sept 2026 read the original ↗
↑