Argo-Bench Pits Data Agents Against a 7.5B-Row Warehouse; Best Model Clears Only 34.8% of Tasks
textql · hf · 2026-10-02
- Design: TextQL's Argo-Bench offers 210 data science/analytics tasks built on a simulated NYC food delivery platform (81M orders in 2024) exported to a 235-table, 7.5-billion-row ERP warehouse modeled on Oracle E-Business Suite. Ground truth is withheld, forcing agents to reconstruct facts by navigating the warehouse.
- Beyond text-to-SQL: Agents take consequential actions (banning fraudulent accounts, allocating courier incentive budgets, issuing back pay) graded by their effects in the simulator; every task has an executable reference solution.
- Results: Across 14 frontier and open-weight models, the best scores 95+ on only 34.8% of tasks and averages 59.5 points, showing enterprise-scale data workflows remain largely unsolved.
More from coding & agent
- OSS contributor slams flood of AI-generated PRs that burden maintainers — Abhishekcur · 2026-10-02
- OpenTag hits #13 on GitHub: open-source AI on-call triage bot for Slack and Teams — FinanceYF5 · 2026-10-02
- Open-source OpenDots chases OpenAI's Dots: self-hosted always-on AI coworkers — FinanceYF5 · 2026-10-02
- Robotic gripper now runs on Opus-written code, aligning parts via built-in light imaging — ihorbeaver · 2026-10-02
- Idempotency vs Deduplication: The Distributed Systems Concepts Engineers Keep Mixing Up — _jaydeepkarale · 2026-10-02
- CodexBar: open-source menu bar app showing AI coding quotas (22k stars) — lxfater · 2026-10-02