TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
The paper presents TraceDance, an agent system that automatically constructs targeted benchmarks for undesirable behaviors from real-world deployment traces. It employs Anchor-and-Confirm for efficient retrieval and confirmation by a Flash LLM, alongside an Anchor Synthesis Loop to revise specifications. Benchmarks evaluate an LLM's next turn using decision-point continuation without requiring reference answers or environment replays. Experiments across hundreds of thousands of sessions show that frontier LLMs struggle significantly on the constructed benchmarks.
Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests.
Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators.
Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points.