CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 13 upvotes

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

QUESTION — How can diverse execution trajectories from specialized agent harnesses be systematically scaled and reconstructed into reusable training data?

The authors propose Recursive Self-Rewrite (RSR), a framework that leverages a base model, Qwen-3.8-27B, to discover successful solutions across diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for leakage and guides revision, and an executor executes them in fresh sandboxes. Supervised finetuning on these expanded trajectories significantly outperforms the base model and direct trajectory SFT across multiple terminal benchmarks.

Three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool.

RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning.

Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2.

taesiri · 02 Oct 2026 read the original ↗
↑