Harness-Zero: Harness Distillation via Agent-as-Harness
The authors study agent harness distillation, using a domain- or instance-optimized harness as training-time guidance to transfer its induced behaviors into model weights. They introduce Harness-Zero, which uses an agent-as-harness to correct student responses before execution in the target action space, turning harness guidance into training demonstrations. Fine-tuning on these trajectories internalizes the harness-induced behavior, allowing the specialized harness to be removed at deployment. Experiments spanning knowledge work, tool use, and science domains show that fine-tuning enables the base model to surpass the performance it reached with the specialized harness attached.
For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness.
With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached.
Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.