RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning often use on-policy distillation (OPD) with a privileged teacher for dense token-level supervision, but privileged info is not always reliable and teacher benefits are stage-dependent. We propose RetireOPD (Self-Retiring On-Policy Distillation), optimizing a decoupled teacher with environment rewards and training a student jointly with RL and OPD. Through Adaptive Retirement, the student drops the teacher once discrepancy stops shrinking and it reaches a target success rate, proceeding with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%.
RetireOPD adopts Adaptive Retirement where the student drops the teacher once discrepancy stops shrinking and reaches a target success rate.
It improves ALFWorld success rate over RL baseline by 14.1% to 18.8% across Qwen2.5 models from 1.5B to 7B.
WebShop accuracy improves by 11.8% to 19.0%, surpassing its own skill-conditioned teacher in every setting.