CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 18 upvotes

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

QUESTION — How can failed LLM agent rollouts be systematically leveraged to improve both error diagnosis and actor recovery?

The authors introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs across diverse environments, along with the Agentic Error-to-Training (AET) pipeline. This pipeline collects natural agent failures, generates corrections, and verifies them against execution logs. Results demonstrate that fine-tuning on these error-diagnosis pairs substantially increases verifier pass rates and improves exact-step agreement in models like Qwen3-8B compared to standard prompting baselines.

First-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points.

Full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout.

Action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.

Leozkl · 30 Sept 2026 read the original ↗
↑