Improving Test-Time Scaling with Adaptive Looped Transformers
L'étude analyse comment les transformateurs bouclés existants gaspillent des itérations sur des jetons inutiles. Les auteurs proposent TaH2, qui post-entraîne conjointement le modèle de base et un décideur d'itération grâce à une supervision de profondeur par anticipation. Sur les benchmarks AIME stimulants, TaH2 améliore considérablement la pente précision-calcul et dépasse la précision maximale du modèle de base.
On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
As the maximum iteration depth instances, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8.