Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
The authors propose Recursive Self-Rewrite (RSR), a framework that leverages a base model, Qwen-3.8-27B, to discover successful solutions across diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for leakage and guides revision, and an executor executes them in fresh sandboxes. Supervised finetuning on these expanded trajectories significantly outperforms the base model and direct trajectory SFT across multiple terminal benchmarks.
Three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool.
RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning.
Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2.