A text-to-SQL paper says benchmark annotation errors can distort rankings and scores
ddkang · x · 2026-07-28
- The paper argues that annotation quality is critical for text-to-SQL benchmarks and warns agent developers not to trust leaderboards blindly.
- It reports recurring annotation-error patterns and says the audit process combines an automated diagnostic agent with human SQL experts.
- The broader claim is that benchmark errors can materially distort both absolute performance and rank ordering.
More from Research
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23