ACLArena: Agent Continue Learning in Multi-stage Post-training
This paper addresses the absence of a well-established recipe for Agent Continual Learning (ACL) across multi-stage post-training by introducing ACLArena, a comprehensive evaluation and analysis framework. The authors examine forgetting and generalization mechanisms at both the model and token levels, systematically comparing multi-teacher on-policy distillation, self-distilled fine-tuning, and model merging. Finally, they propose a novel ACL recipe combining offline replay over high-quality trajectories with a routed network of multiple LoRA experts specialized via reinforcement learning.
A sequential training pipeline explains the mechanisms of forgetting and generalization from model-level and token-level perspectives.
A new ACL recipe combining offline replay over high-quality trajectories with a routed network of multiple LoRA experts specialized via RL substantially improves multi-domain learning.