CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 55 upvotes

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

QUESTION — How can multi-turn agents be trained using reinforcement learning while avoiding instability from privileged teacher information?

Multi-turn agents trained with reinforcement learning often use on-policy distillation (OPD) with a privileged teacher for dense token-level supervision, but privileged info is not always reliable and teacher benefits are stage-dependent. We propose RetireOPD (Self-Retiring On-Policy Distillation), optimizing a decoupled teacher with environment rewards and training a student jointly with RL and OPD. Through Adaptive Retirement, the student drops the teacher once discrepancy stops shrinking and it reaches a target success rate, proceeding with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%.

RetireOPD adopts Adaptive Retirement where the student drops the teacher once discrepancy stops shrinking and reaches a target success rate.

It improves ALFWorld success rate over RL baseline by 14.1% to 18.8% across Qwen2.5 models from 1.5B to 7B.

WebShop accuracy improves by 11.8% to 19.0%, surpassing its own skill-conditioned teacher in every setting.

LZXzju · 17 Sept 2026 read the original ↗
↑