ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
QUESTION — How can historical self-generated rollouts be reused to improve LLM post-training without requiring additional policy rollouts?
The authors introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), a method that preserves full historical trajectories as context while applying loss only to selected model-generated continuations. This addresses the issue of stale historical data or data containing mistakes. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results demonstrate that self-rollout training leaves behind reusable behavioral experience exploitable via offline SFT.
On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20%.
ROSS improves SWE-bench Verified from 64.20% to 68.40%.
Hiiamein · 28 Sept 2026
read the original ↗