CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 10 upvotes

On the Off-Policy Teacher in On-Policy Distillation

QUESTION — How can the performance degradation of the teacher when supervising student-generated prefixes in on-policy distillation be mitigated?

This paper addresses the asymmetry in on-policy distillation (OPD) where the teacher must supervise student-generated prefixes, which causes its continuation performance to degrade. To fix this, the authors propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's continuation capability and consistently improves on-policy distillation effectiveness across multiple configurations.

SCOUT improves the teacher's ability to continue from student-generated prefixes.

SCOUT co-trains the teacher by periodically optimizing its conditional ability using reinforcement learning with verifiable rewards alongside standard OPD updates.

SCOUT consistently improves the effectiveness of on-policy distillation across multiple teacher-student configurations, model scales, and reasoning domains.

shrango · 29 Sept 2026 read the original ↗
↑