CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 158 upvotes

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

QUESTION — How can the long-horizon decision-making capability (taste) of LLM agents be measured and improved in engineering and research tasks?

This research defines an agent's 'taste' as the ability to make correct long-horizon decisions at critical forks during engineering and research tasks. The authors build Taste-Bench, a benchmark that automatically mines decision forks from parallel attempts and detours without human annotation. Evaluations show that the best model answers only 59.7% of the questions correctly, and increasing the reasoning budget does not improve accuracy. By distilling the judgment of a teacher model into a student model, the approach improves decision-making on unseen tasks and enhances end-to-end success on held-out SWE-bench Pro tasks.

The best model answers only 59.7% of the questions correctly on Taste-Bench.

Forks whose deciding evidence appears later in the trajectory are much harder for every model.

A larger reasoning budget does not improve the accuracy on Taste-Bench.

wenbopan · 22 Sept 2026 read the original ↗
↑