Abstract
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
Community
Hi everyone, first author here!
We realized most established text-to-SQL benchmarks don't yet cover the most important care most companies care about: the agents' ability to make good decisions.
While some benchmarks have been pretty good for this at a micro-level (e.g., Vending-Bench 2 by Andon Labs), none have had the ambition to test how well LLMs would do making day-to-day decisions across different areas, like forecasting or fraud prevention, at very large companies.
We've gotten to some extremely interesting findings on what LLMs struggle with here, and are really excited to share our findings in the Argo-Bench preprint.
A demo of the world is available on our project page – check it out!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline (2026)
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents (2026)
- Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge (2026)
- Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents (2026)
- Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses (2026)
- APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows (2026)
- CausalVerify: End-to-End Verification of Causal Analyses by Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
textql/Decision-Bench
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper