MInTRL: Off-policy Intervention can boost On-policy RL
Reinforcement learning with verifiable rewards is typically on-policy, limiting learning to self-discovered trajectories, whereas off-policy methods leverage external knowledge but suffer from distribution shifts. This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands exploration via sparse, local interventions in otherwise on-policy rollouts using a judge-intervention policy to replace erroneous suffixes with corrections. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines.
Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines.