Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
This work addresses the limitation of finite rollout budgets in Reinforcement Learning with Verifiable Rewards (RLVR), which frequently produce all-fail groups lacking policy-gradient signals. The authors propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups from heterogeneous models. By controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping, GRAFT enables mutual learning without a designated stronger teacher, consistently improving model performance across mathematical reasoning benchmarks.
Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance.
Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.