Improving Test-Time Scaling with Adaptive Looped Transformers
The study analyzes how existing looped transformers often yield steep accuracy-compute slopes but underperform baselines at matched compute because fixed-depth looping wastes iterations on unhelpful tokens. The authors propose TaH2, which jointly post-trains the backbone and an iteration decider through lookahead depth supervision using online labels. On challenging AIME benchmarks, TaH2 significantly improves the accuracy-compute slope and exceeds the baseline's peak accuracy.
On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8.