Corrected text-to-SQL data shifts 16 open-source agent rankings by up to nine places
ddkang · x · 2026-07-28
- Re-evaluating 16 open-source text-to-SQL agents on corrected benchmark data changes rankings by as much as -9 to +9.
- Former SOTA Contextual-SQL drops from #1 to #7, while GenaSQL and CHESS move up to #1.
- The chart compares original and corrected execution accuracy to show how sensitive leaderboard positions are to annotation quality.
More from Research
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23