BiomedSQL: New Biomedical Text-to-SQL Benchmark Shows Top Models Far Below Expert Level
DanielKhashabi · x · 2026-09-30
Researchers from Johns Hopkins released BiomedSQL on arXiv, the first benchmark explicitly evaluating scientific reasoning in text-to-SQL generation over a real-world biomedical knowledge base.
- Scale: 68,000 question/SQL/answer triples grounded in a harmonized BigQuery database integrating gene-disease associations, omics-based causal inference, and drug approval records
- Focus: questions require inferring implicit domain criteria—genome-wide significance thresholds, effect directionality, trial phase filtering—rather than syntactic translation alone
- Large gap: Gemini-3-Pro reaches 58.1% execution accuracy with baseline prompting; the team's multi-step BMSQL agent hits 62.6%, both well below the 90.0% expert baseline
The benchmark sets a new foundation for measuring LLM scientific reasoning.
More from Research
- DexAgent learns bimanual dexterous manipulation from one human video, hitting 63.6% success — YunzhuLiYZ · 2026-09-30
- DexTaG uses human tactile demos as RL reward guidance to teach robots human-like tool use — gan_chuang · 2026-09-30
- AnyBokeh: Editing Bokeh Directly in Blurred Photos, No Sharp Image Needed — ccloy · 2026-09-30
- DynaRobotics trains robust whole-body controller for wheeled robot via Isaac Sim RL — zhengyiluo · 2026-09-30
- Google: selecting diverse reasoning routes in SFT improves post-RL generalization by up to 16.9 points — google · 2026-09-30
- Rowmax-H15: approximate softmax in attention for 25.8% faster B200 inference with minimal quality loss — illinois · 2026-09-30