On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
This work evaluates the roles of rollout policy, token-level KL direction, and learning rate in strong-to-weak model distillation across the Llama3 and Qwen2.5 model families on scientific, medical, and arithmetic tasks. By independently varying these factors, the authors discover that rollout policy does not inherently play a central role as commonly assumed. Instead, token-level KL direction dictates task performance and output coverage, while learning rate governs forgetting and update sparsity. Forward KL remains remarkably robust to rollout policies, whereas reverse KL is much more sensitive and favors student-generated rollouts.
Forward KL exhibits exceptional robustness to rollout policy changes.
Reverse KL displays significantly higher sensitivity and prefers student-generated rollouts.
On-policy data improves generalisation to harder variants of the Countdown arithmetic task.