CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 144 upvotes

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

QUESTION — How can one overcome the instability in output-space extrapolation during on-policy distillation?

This paper addresses the instability and performance degradation in output-space extrapolation during on-policy distillation (OPD). The authors introduce RIDE (RL-Induced Direction Extrapolation), a method that extrapolates the RL-induced change directly in representation space rather than output space. By computing the residual between the teacher and its pre-RL checkpoint across every layer and token position, RIDE regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Experiments across four base/RL-teacher pairs demonstrate that RIDE consistently approaches or exceeds the RL-trained teacher's performance.

RIDE extrapolates the RL-induced change directly in representation space across every layer and token position.

The regression is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher.

RIDE approaches or exceeds the RL-trained teacher on every tested model pair.

LH2101 · 29 Sept 2026 read the original ↗
↑