Study finds BIRD and Spider 2.0-Snow text-to-SQL benchmarks exceed 50% annotation error
ddkang · x · 2026-07-28
Key finding
The authors report that two widely used text-to-SQL benchmarks, BIRD and Spider 2.0-Snow, contain pervasive annotation errors. In their VLDB 2026 paper, they say expert review found error rates above 50%, and that these mistakes can shift measured agent performance by up to 19%.
Impact
The paper argues that benchmark leaderboards for text-to-SQL are less reliable than they appear, because annotation quality directly affects both evaluation and ranking.
More from coding & agent
- Study of 100,000 developers finds AI coding gains shrink to about 30% at release stage — amcafee · 2026-07-28
- A new guide says most codebases are not ready for cloud coding agents — vinvan · 2026-07-28
- Meme asks whether to write code now or wait for a model that one-shots it — Darpinian · 2026-07-28
- UWaterloo open-sources Interactive Training 2 for auditable live model training — UWaterloo · 2026-07-28
- LangChain highlights dcode for swapping GPT-5.6 to Kimi K3 in 10 seconds — LangChain · 2026-07-28
- Waddle Labs pitches "Claude Code for robots" with 20-minute task execution — ycombinator · 2026-07-28