First Text-to-SQL Model Beats Human Benchmark Using RLVR
EchoShao8899 · x · 2026-08-28
Researchers from UIUC and Bridgewater, in collaboration with Thinking Machines Lab, used Reinforcement Learning from Verifiable Rewards (RLVR) with expert-aligned data cleaning. This approach produced the first text-to-SQL model to surpass human performance (92.96%) on the BIRD benchmark. While frontier models like GPT-5.6 score in the mid-80s, they struggle with ambiguity and high costs.
Related event: RLVR with Cleaned Data First Beats Human Baseline on Text-to-SQL(6 posts)→
More from Research
- Co-Scientist evaluation: Severe hallucinations drop to 4%, fabrication to 0% — SRSchmidgall · 2026-08-29
- AI Scientist limitations: Selective reporting and code-method divergence remain — SRSchmidgall · 2026-08-29
- Gemini Designs Precursor for MXene Synthesis in New Materials Science Experiment — SRSchmidgall · 2026-08-29
- Co-Scientist deployed across materials science, biology, and CS with varying autonomy — SRSchmidgall · 2026-08-29
- Google releases Co-Scientist to accelerate real-world scientific discovery — SRSchmidgall · 2026-08-29
- CarNet Matches Spherical-Harmonics Accuracy in Cartesian Space for Atomic Simulations — bravo_abad · 2026-08-29