Study finds BIRD and Spider 2.0-Snow text-to-SQL benchmarks exceed 50% annotation error

ddkang · x · 2026-07-28

Key finding

The authors report that two widely used text-to-SQL benchmarks, BIRD and Spider 2.0-Snow, contain pervasive annotation errors. In their VLDB 2026 paper, they say expert review found error rates above 50%, and that these mistakes can shift measured agent performance by up to 19%.

Impact

The paper argues that benchmark leaderboards for text-to-SQL are less reliable than they appear, because annotation quality directly affects both evaluation and ranking.

Related event: VLDB Paper Reveals Over 50% Mislabeled Rates in Major Text-to-SQL Benchmarks(7 posts)→

Original post →

More from coding & agent

coding & agent channel →