Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks simulating a food delivery platform in New York City at true scale, with 81 million orders in 2024 exported to an ERP warehouse of 235 tables and 7.5 billion rows. The simulator's ground-truth state is withheld from the warehouse the agent sees, requiring tasks to reconstruct facts by navigating the warehouse before filing actions such as banning fraudulent accounts or allocating incentive budgets. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.
Argo-Bench simulates a food delivery platform in New York City with 81 million orders in 2024.
The ERP warehouse contains 235 tables and 7.5 billion rows modeled on the Oracle E-Business Suite schema.
The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.