CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 26 upvotes

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

QUESTION — How can data agents and analytics workflows be evaluated on enterprise-scale data warehouses that require navigating complex schemas and executing consequential actions?

We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks simulating a food delivery platform in New York City at true scale, with 81 million orders in 2024 exported to an ERP warehouse of 235 tables and 7.5 billion rows. The simulator's ground-truth state is withheld from the warehouse the agent sees, requiring tasks to reconstruct facts by navigating the warehouse before filing actions such as banning fraudulent accounts or allocating incentive budgets. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.

Argo-Bench simulates a food delivery platform in New York City with 81 million orders in 2024.

The ERP warehouse contains 235 tables and 7.5 billion rows modeled on the Oracle E-Business Suite schema.

The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.

gtomitsu · 01 Oct 2026 read the original ↗
↑