Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Terminal agents execute stochastic model generations, but flawed commands can disrupt the environment and hinder progress. The authors investigate Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before execution while keeping the generator intact. Experiments on TerminalBench-Lite show that employing a capable verifier, such as GPT-5.6 Sol, allows the agent to exploit useful alternative actions from the generator, substantially raising Pass@1 accuracy. Distilling responses from a stronger verifier further boosts performance at lower token costs than raw trajectory scaling.
On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions.
Combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone.