Study finds BIRD and Spider 2.0-Snow text-to-SQL benchmarks exceed 50% annotation error
ddkang · x · 2026-07-28
Key finding
The authors report that two widely used text-to-SQL benchmarks, BIRD and Spider 2.0-Snow, contain pervasive annotation errors. In their VLDB 2026 paper, they say expert review found error rates above 50%, and that these mistakes can shift measured agent performance by up to 19%.
Impact
The paper argues that benchmark leaderboards for text-to-SQL are less reliable than they appear, because annotation quality directly affects both evaluation and ranking.
More from coding & agent
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Vite+ Hits RC: One Rust-Powered CLI to Replace Your Entire Web Toolchain — cnakazawa · 2026-09-23
- Tesla's in-car Grok agent books trips across Gmail, Calendar and Notion in one command — xiaohu · 2026-09-23
- Tesla's In-Car Grok Assistant Now Executes Cross-App Tasks in One Sentence — xiaohu · 2026-09-23
- Garry Tan says Capy lets him ship PRs much faster than Codex or Claude Code — garrytan · 2026-09-23
- DeskPilot: open-source native Python desktop client for local LLMs with MCP and sandboxed tools — poofph · 2026-09-23