ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
The authors introduce ActiveSaddler, framing automated harness optimization as an automated curriculum learning problem. ActiveSaddler models the evolving training curriculum as a non-stationary bandit with dynamically instantiated optimization targets, abstracting recurring failures into reusable arms that balance revisiting known weaknesses with exploring unseen scenarios. Experiments on GAIA2 and Terminal-Bench 2.0 demonstrate that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over a fixed scenario order, respectively.
ActiveSaddler improves test Pass@1 by 4.4 percentage points on GAIA2 over a harness optimizer using a fixed scenario order.
It improves test Pass@1 by 7.5 percentage points on Terminal-Bench 2.0 over the fixed scenario order baseline.