Two widely used text-to-SQL benchmarks are reported to have 52.8% and 62.8% error rates
ddkang · x · 2026-07-28
- The audit claims BIRD Mini-Dev and Spider 2.0-Snow have annotation error rates of 52.8% and 62.8%.
- The team used a three-stage process: automated diagnostic reports, expert adjudication of flagged cases, and additional review of recurring error patterns.
- It argues these benchmark issues can mislead both research directions and deployment choices.
More from Research
- A deep dive on building frontier-lab evals explains why 100% scores can be a failure — aakashgupta · 2026-07-28
- Maker shares first AI robot kit built with Raspberry Pi 5 and Hermes agent — petrusenko_max · 2026-07-28
- Exploring Artificial Life: Wolfram and Others Feature in Lenia Simulation — max_romana · 2026-07-28
- A new artificial-life video asks what’s missing for open-ended evolution — max_romana · 2026-07-28
- Macrocosmos starts a permissionless 16B model training run across three continents — markjeffrey · 2026-07-28
- Free AI curriculum maps a practical path from first principles to LLMs — tetsuoai · 2026-07-28