CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 12 upvotes

MInTRL: Off-policy Intervention can boost On-policy RL

QUESTION — How can sparse off-policy interventions be integrated into on-policy reinforcement learning to expand exploration without causing large distribution shifts?

Reinforcement learning with verifiable rewards is typically on-policy, limiting learning to self-discovered trajectories, whereas off-policy methods leverage external knowledge but suffer from distribution shifts. This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands exploration via sparse, local interventions in otherwise on-policy rollouts using a judge-intervention policy to replace erroneous suffixes with corrections. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines.

Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines.

MYC081 · 11 Sept 2026 read the original ↗
↑