CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 15 upvotes

LastOPD: Taming Collapse in Latent On-Policy Distillation

QUESTION — How can collapse be overcome during latent on-policy distillation?

On-policy distillation with latent supervision aligns a student's latent states with a teacher's, but frequently suffers from early gains followed by severe collapse in performance despite improving alignment metrics. This occurs due to a mismatch in how latent signals are applied across layers. To address this, the authors propose LastOPD, which applies the latent signal exclusively at the last-layer state—the common interface read by both language model heads—during a 10-step crossfade into token-level OPD. Extensive experiments show that LastOPD improves MATH-500 performance over token-only OPD and achieves final scores in about half the steps.

Latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery.

LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers.

LastOPD reaches the final score of token-only OPD in about half the steps.

Muyiiiii · 23 Sept 2026 read the original ↗
↑