CONSONANCE.for your information
lundi 5 octobre 2026frenvi

À lire de près

01 — evaluation 47 votes

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

QUESTION — Comment construire automatiquement des benchmarks de comportement d'agents à partir de traces de déploiement réelles ?

L'article présente TraceDance, un système d'agents qui construit automatiquement des benchmarks ciblés pour des comportements indésirables à partir de traces de déploiement réelles. Il utilise Anchor-and-Confirm pour une récupération efficace et une confirmation par un Flash LLM, ainsi qu'une boucle de synthèse. Les benchmarks évaluent le tour suivant d'un LLM sans nécessiter de réponses de référence. Les expériences montrent que les LLM actuels peinent à répondre de manière appropriée.

Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests.

Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators.

Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points.

ZhishanQ · 27 sept. 2026 lire l'original ↗
↑