CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 13 upvotes

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

QUESTION — How can historical self-generated rollouts be reused to improve LLM post-training without requiring additional policy rollouts?

The authors introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), a method that preserves full historical trajectories as context while applying loss only to selected model-generated continuations. This addresses the issue of stale historical data or data containing mistakes. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results demonstrate that self-rollout training leaves behind reusable behavioral experience exploitable via offline SFT.

On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20%.

ROSS improves SWE-bench Verified from 64.20% to 68.40%.

Hiiamein · 28 Sept 2026 read the original ↗
↑