CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 221 upvotes

Scaling Properties of Same-Family On-Policy Distillation

QUESTION — What are the scaling properties of on-policy distillation across weak-to-strong, same-base, and strong-to-weak setups?

This paper investigates the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak setups. The authors discover that early OPD training dynamics exhibit a useful-transfer regime where held-out accuracy rises linearly with the square root of token-level reverse KL divergence. Remarkably, in every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own. By fitting power laws for how peak gold score and transfer slope scale with parameter counts and teacher scores, the work demonstrates that compact experts can effectively transfer capabilities and that smaller teachers with matched scores can transfer better.

Held-out accuracy rises approximately linearly in the square root of token-level reverse KL divergence during the useful-transfer regime.

In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own.

Peak gold score improves with teacher scale only up to roughly the student's scale.

colored-dye · 26 Sept 2026 read the original ↗
↑