What Makes Recurrence Effective in Looped Language Models?
This paper investigates looped language models (LoopLMs), which increase computational depth through parameter sharing without adding parameters. Through controlled experiments, the authors systematically examine when recurrence helps, where it should be applied, and how conditioning affects performance across budgets. They find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance. To address limitations of initial-state injection, they propose history-state injection combined with timestep conditioning as a more effective design that better preserves knowledge under extended unrolling.
Recurrence can improve reasoning beyond the training horizon while degrading knowledge performance.
Non-recurrent output layers improve robustness to under-unrolling, while preferred layer placement varies with inference budget.
Channel-wise history-state injection combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness.