Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
This study addresses automated root-cause attribution (RCA) for failures in long-horizon AI agents, where execution logs are massive and relevant information is sparse. The authors introduce Continual Search, an iterative framework that nudges the LLM judge over successive turns to keep searching for unresolved diagnostic evidence, avoiding premature conclusions. To evaluate RCA at scale, they introduce MegaRCA-Mix, consisting of 50 human-annotated failure trials across long-horizon tasks. Experiments show significant performance gains, demonstrating that effective search supersedes raw model scale.
MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks.
On MegaRCA-Mix, it improves GPT-5.5's F1 score by more than 40%, from 0.349 to 0.498.
Within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.