CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 10 upvotes

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

QUESTION — How can sample efficiency and rollout efficiency in on-policy distillation for LLMs be improved using reinforcement learning?

This paper studies on-policy distillation (OPD) through the lens of reinforcement learning, connecting the reverse-KL objective to KL-regularized policy optimization. The authors introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that integrates optimistic exploration and off-policy data reuse into policy distillation. LSPD preserves policy diversity and improves rollout efficiency by repeatedly learning from previously collected trajectories. Empirical evaluations across six mathematical reasoning benchmarks show that LSPD achieves average gains of +1.59 points in Avg@16, and its fully off-policy variant matches vanilla OPD performance using only the first 25% of rollout batches.

LSPD achieves average gains of +1.59 points in Avg@16 across six mathematical reasoning benchmarks.

The idealized formulation of LSPD achieves a sharp mathcal O(log K) regret bound under online exploration.

Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches.

DVA13304 · 28 Sept 2026 read the original ↗
↑