CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 13 upvotes

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

QUESTION — How can teacher-side self-deviations be eliminated during on-policy distillation to improve reasoning model training?

This paper demonstrates that the token-level discrepancy between a teacher and an on-policy student in standard on-policy distillation (OPD) contains deviations originating from the teacher itself, which are incorrectly learned by the student. To fix this, the authors introduce Calibrated On-Policy Distillation (Cal-OPD), which estimates the teacher's self-deviation region using positive and negative privileged interventions and retains only the component beyond this region. Experiments on mathematical reasoning benchmarks show that Cal-OPD consistently outperforms standard OPD across various model scales.

Retains only about 52--65% of the original teacher--student discrepancy as the optimization signal.

Consistently outperforms standard OPD and its variants across model scales.

Evaluated through experiments on mathematical reasoning benchmarks.

qqiang · 18 Sept 2026 read the original ↗
↑