LastOPD: Taming Collapse in Latent On-Policy Distillation
On-policy distillation with latent supervision aligns a student's latent states with a teacher's, but frequently suffers from early gains followed by severe collapse in performance despite improving alignment metrics. This occurs due to a mismatch in how latent signals are applied across layers. To address this, the authors propose LastOPD, which applies the latent signal exclusively at the last-layer state—the common interface read by both language model heads—during a 10-step crossfade into token-level OPD. Extensive experiments show that LastOPD improves MATH-500 performance over token-only OPD and achieves final scores in about half the steps.
Latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery.
LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers.
LastOPD reaches the final score of token-only OPD in about half the steps.