The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
This paper addresses the instability and performance degradation in output-space extrapolation during on-policy distillation (OPD). The authors introduce RIDE (RL-Induced Direction Extrapolation), a method that extrapolates the RL-induced change directly in representation space rather than output space. By computing the residual between the teacher and its pre-RL checkpoint across every layer and token position, RIDE regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Experiments across four base/RL-teacher pairs demonstrate that RIDE consistently approaches or exceeds the RL-trained teacher's performance.
RIDE extrapolates the RL-induced change directly in representation space across every layer and token position.
The regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.
RIDE approaches or exceeds the RL-trained teacher on every tested model pair.