AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
QUESTION — How can an AI agent's ability to understand, manipulate, and improve training data be isolated and systematically evaluated?
The paper introduces AutoDataBench, a controlled testbed to evaluate the Data Intelligence of AI agents. The framework focuses on data diagnosis, organization, and construction through three curated optimization tasks while holding non-data factors constant. It tests frontier LLMs' capacity to improve training data via iterative experimentation under specific resource budgets, and demonstrates that reusing trajectories improves downstream coding performance during mid-training.
gowitheflow · 30 Sept 2026
read the original ↗