TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
Time series agents often struggle with fixed tool libraries curated in advance by humans, leading to reduced accuracy and flawed self-revision cycles. To address this, the authors propose TimeEvo, a failure-driven self-evolution framework. TimeEvo clusters diagnosed agent failures into capability gaps, plans measurements, synthesizes evidence-only tools, and admits candidate libraries via a paired admission gate. Experiments across ten time series QA tasks and three backbones demonstrate that TimeEvo improves accuracy on every task and backbone tested.
A library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test.
One round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point.
Starting from an empty library, TimeEvo improves accuracy on every task and every backbone.