On the Off-Policy Teacher in On-Policy Distillation
This paper addresses the asymmetry in on-policy distillation (OPD) where the teacher must supervise student-generated prefixes, which causes its continuation performance to degrade. To fix this, the authors propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's continuation capability and consistently improves on-policy distillation effectiveness across multiple configurations.
SCOUT improves the teacher's ability to continue from student-generated prefixes.
SCOUT co-trains the teacher by periodically optimizing its conditional ability using reinforcement learning with verifiable rewards alongside standard OPD updates.
SCOUT consistently improves the effectiveness of on-policy distillation across multiple teacher-student configurations, model scales, and reasoning domains.