An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
This paper studies on-policy distillation (OPD) through the lens of reinforcement learning, connecting the reverse-KL objective to KL-regularized policy optimization. The authors introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that integrates optimistic exploration and off-policy data reuse into policy distillation. LSPD preserves policy diversity and improves rollout efficiency by repeatedly learning from previously collected trajectories. Empirical evaluations across six mathematical reasoning benchmarks show that LSPD achieves average gains of +1.59 points in Avg@16, and its fully off-policy variant matches vanilla OPD performance using only the first 25% of rollout batches.
LSPD achieves average gains of +1.59 points in Avg@16 across six mathematical reasoning benchmarks.
The idealized formulation of LSPD achieves a sharp mathcal O(log K) regret bound under online exploration.
Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches.