First Text-to-SQL Model Beats Human Benchmark Using RLVR

EchoShao8899 · x · 2026-08-28

Researchers from UIUC and Bridgewater, in collaboration with Thinking Machines Lab, used Reinforcement Learning from Verifiable Rewards (RLVR) with expert-aligned data cleaning. This approach produced the first text-to-SQL model to surpass human performance (92.96%) on the BIRD benchmark. While frontier models like GPT-5.6 score in the mid-80s, they struggle with ambiguity and high costs.

Related event: RLVR with Cleaned Data First Beats Human Baseline on Text-to-SQL(6 posts)→

Original post →

More from Research

Research channel →