ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
The paper introduces ModularRSI, a benchmark-disjoint, contrastive, and modular framework designed to evolve agent harnesses reliably. It decomposes harnesses into five functional modules—such as Agent Loop and Tool Use—and evaluates them using 2,000 external evolution tasks. By contrasting successful and failed trajectories across tasks, ModularRSI isolates recurring behavioral deficiencies without overfitting to evaluation benchmarks. Experiments on TB2.0 and SWE-Bench Verified demonstrate consistent improvements on unseen tasks and robust transferability across different foundation models.
ModularRSI curates 2,000 executable evolution tasks from external sources that are disjoint from downstream evaluation benchmarks.
The system decomposes the evolvable harness into five functional modules that evolve independently within restricted scopes.
Experiments on TB2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks.