Two widely used text-to-SQL benchmarks are reported to have 52.8% and 62.8% error rates
ddkang · x · 2026-07-28
- The audit claims BIRD Mini-Dev and Spider 2.0-Snow have annotation error rates of 52.8% and 62.8%.
- The team used a three-stage process: automated diagnostic reports, expert adjudication of flagged cases, and additional review of recurring error patterns.
- It argues these benchmark issues can mislead both research directions and deployment choices.
More from Research
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23