Calibrating Teacher--Student Discrepancy for On-Policy Distillation
This paper demonstrates that the token-level discrepancy between a teacher and an on-policy student in standard on-policy distillation (OPD) contains deviations originating from the teacher itself, which are incorrectly learned by the student. To fix this, the authors introduce Calibrated On-Policy Distillation (Cal-OPD), which estimates the teacher's self-deviation region using positive and negative privileged interventions and retains only the component beyond this region. Experiments on mathematical reasoning benchmarks show that Cal-OPD consistently outperforms standard OPD across various model scales.
Retains only about 52--65% of the original teacher--student discrepancy as the optimization signal.
Consistently outperforms standard OPD and its variants across model scales.
Evaluated through experiments on mathematical reasoning benchmarks.