Corrected text-to-SQL data shifts 16 open-source agent rankings by up to nine places
ddkang · x · 2026-07-28
- Re-evaluating 16 open-source text-to-SQL agents on corrected benchmark data changes rankings by as much as -9 to +9.
- Former SOTA Contextual-SQL drops from #1 to #7, while GenaSQL and CHESS move up to #1.
- The chart compares original and corrected execution accuracy to show how sensitive leaderboard positions are to annotation quality.
More from Research
- Kimi K3 report details kernel tuning, a Triton-like compiler, and a chip prototype — stochasticchasm · 2026-07-28
- Follow-up on the artificial-life video adds a Lenia simulator link — max_romana · 2026-07-28
- A new artificial-life video asks what’s missing for open-ended evolution — max_romana · 2026-07-28
- Macrocosmos starts a permissionless 16B model training run across three continents — markjeffrey · 2026-07-28
- Free AI curriculum maps a practical path from first principles to LLMs — tetsuoai · 2026-07-28
- Kimi report reveals a wide internal benchmark suite for coding and agent skills — stochasticchasm · 2026-07-28