The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
This research defines an agent's 'taste' as the ability to make correct long-horizon decisions at critical forks during engineering and research tasks. The authors build Taste-Bench, a benchmark that automatically mines decision forks from parallel attempts and detours without human annotation. Evaluations show that the best model answers only 59.7% of the questions correctly, and increasing the reasoning budget does not improve accuracy. By distilling the judgment of a teacher model into a student model, the approach improves decision-making on unseen tasks and enhances end-to-end success on held-out SWE-bench Pro tasks.
The best model answers only 59.7% of the questions correctly on Taste-Bench.
Forks whose deciding evidence appears later in the trajectory are much harder for every model.
A larger reasoning budget does not improve the accuracy on Taste-Bench.