Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
The authors introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs across diverse environments, along with the Agentic Error-to-Training (AET) pipeline. This pipeline collects natural agent failures, generates corrections, and verifies them against execution logs. Results demonstrate that fine-tuning on these error-diagnosis pairs substantially increases verifier pass rates and improves exact-step agreement in models like Qwen3-8B compared to standard prompting baselines.
First-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points.
Full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout.
Action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.