Human-Level Text-to-SQL via RLVR and Verified Data Cleaning
ddkang · x · 2026-08-28
Paper proposes achieving human-level Text-to-SQL by fine-tuning Qwen3-235B using RLVR (Reinforcement Learning on Verified Data) on clean data, without complex pipelines.
Key findings:
- Existing training data contains pervasive annotation errors that mislead optimization.
- Developed a multi-round expert verification pipeline to curate BIRD-Platinum (2.5k instances from BIRD Train), correcting errors in 61% of instances.
- Fine-tuning on BIRD-Platinum yields 11-16% improvements over BIRD Train on Arcwise-Plat and Spider2, outperforming SOTA open-source systems.
More from Research
- OpenResearch Launches AutoResearch: Automating Paper Replication with Agent Swarms — simonguozirui · 2026-08-28
- BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark — himanshustwts · 2026-08-28
- AI Methods Match Humans in Designing RNA Structures, Paving Way for New Therapeutics — rishabh16_ · 2026-08-28
- Auto-research loops are the future: RL discovers novel states 5x more efficiently — const_reborn · 2026-08-28
- Intrinsic Discovery method enables unsupervised model exploration — burny_tech · 2026-08-28
- Active learning scans 2.5M molecules to find new way to block RAS-driven cancers — bravo_abad · 2026-08-28