CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 177 upvotes

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

QUESTION — How do rollout policy, token-level KL direction, and learning rate affect model distillation dynamics?

This work evaluates the roles of rollout policy, token-level KL direction, and learning rate in strong-to-weak model distillation across the Llama3 and Qwen2.5 model families on scientific, medical, and arithmetic tasks. By independently varying these factors, the authors discover that rollout policy does not inherently play a central role as commonly assumed. Instead, token-level KL direction dictates task performance and output coverage, while learning rate governs forgetting and update sparsity. Forward KL remains remarkably robust to rollout policies, whereas reverse KL is much more sensitive and favors student-generated rollouts.

Forward KL exhibits exceptional robustness to rollout policy changes.

Reverse KL displays significantly higher sensitivity and prefers student-generated rollouts.

On-policy data improves generalisation to harder variants of the Countdown arithmetic task.

jpiskorz · 28 Sept 2026 read the original ↗
↑