CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 24 upvotes

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

QUESTION — How can we leverage successful trajectories from heterogeneous peer models to overcome all-fail groups in RLVR?

This work addresses the limitation of finite rollout budgets in Reinforcement Learning with Verifiable Rewards (RLVR), which frequently produce all-fail groups lacking policy-gradient signals. The authors propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups from heterogeneous models. By controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping, GRAFT enables mutual learning without a designated stronger teacher, consistently improving model performance across mathematical reasoning benchmarks.

Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance.

Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.

jadohu · 29 Sept 2026 read the original ↗
↑