New Data Agent Benchmark: Best Frontier Agent Passes Just a Third of 54 Multi-DB Queries

CShorten30 · x · 2026-09-14

Shreya Shankar (UC Berkeley) surveyed 25+ data agent benchmarks and found none testing queries across multiple databases. Her team's new Data Agent Benchmark spans 54 queries, 12 datasets, and 4 database systems — and the best frontier agent passes barely a third.

Failure modes expose intrinsic LLM biases: facing unfamiliar SQL dialects, agents dodge by dumping tables to flat files and doing everything in Pandas.

Full conversation on Weaviate Podcast #135.

Original post →

More from coding & agent

coding & agent channel →