BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
Mathew J. Koretsky, Maya Willey, Owen Bianchi, Chelsea X. Alvarado, Tanay Nayak, Nicole Kuznetsov, Sungwon Kim, Mike A. Nalls, Daniel Khashabi, Faraz Faghri
COLM 2026
cs.CL, cs.AI, cs.LG
2025-05-24
BiomedSQL holds 68k biomedical text-to-SQL items needing implicit scientific reasoning. Gemini-3-Pro hits 58.1% execution accuracy, BMSQL 62.6%, vs a 90% expert baseline.
Spider and BIRD score whether a model can map English onto SQL across many schemas. MIMICSQL and EHRSQL score patient retrieval and temporal logic over EHRs. Biomedical analysts ask a different class of question. Which SNPs are most significantly associated with Alzheimer's, or which drugs target genes up-regulated in Parkinson's, only become executable SQL after someone fills in conventions the schema never states: genome-wide significance at p < 5×10⁻⁸, effect direction from the beta coefficient, approval as an indication-specific trial-phase filter.
BiomedSQL is built for that gap. NIH CARD, DataTecnica, and Johns Hopkins released it at COLM 2026. The task is narrow: take a qualitative scientific question over a real biomedical warehouse and emit SQL that actually runs.
The warehouse is a ten-table BigQuery database, with tables up to 72 million rows and 31 columns. BigQuery is a deliberate dialect choice: cloud-native SQL is common in biomedical pipelines, and vendor functions are under-tested. It joins Open Targets gene-disease-drug links, ChEMBL pharmacology, SNP-level GWAS summary statistics for Alzheimer's and Parkinson's from the GWAS Catalog, and omicSynth causal estimates from summary-data-based Mendelian randomization (SMR), which treats genetic variants as instruments for causal inference.
Forty seed questions from CARDBiomedBench were written as gold SQL by one domain expert. Two analysts independently checked execution output and a natural-language answer; disagreements were discussed to a single consensus. Separate annotations were not kept, so there is no inter-annotator agreement figure. Substituting disease, gene, and SNP entities scaled the set to 68,227 question/SQL/answer triples. Gold queries avoid SELECT and cap at 100 rows. Mean SQL length is 96.4 tokens, longer than BIRD's 50.6, second only to EHRSQL.
The test set is 546 questions sampled to match the template mix. SQL is scored three ways. Execution Accuracy (EX) is 1 only when the result set matches gold exactly (UUID sets for SELECT, numeric values for COUNT and similar). Jaccard (JAC) credits partial overlap. Syntax Error Rate (SER) is the share of queries that do not run. Natural-language answers go through BioScore on GPT-4o: Response Quality Rate (RQR) for factual correctness, Safety Rate (SR) for abstention among wrong or refused answers. On 100 double-graded items, GPT-4o and a domain expert correlated at Spearman 0.89.
Models: Llama-3.1 70B and 405B, Qwen-2.5-Coder 14B and 32B, GPT-4o, GPT-o3-mini, GPT-5.2, Gemini-2.0-Flash, Gemini-3-Pro, Claude-3.7-Sonnet, Claude-4.5-Opus. Prompts add sample rows, few-shot examples, and explicit statistical-threshold instructions. Interaction setups include ReAct, LlamaIndex schema retrieval, an adapted DAIL-SQL, and BMSQL, a custom loop that rewrites a query from intermediate results or execution errors the way an analyst debugs SQL.
The expert baseline is the mean of two biomedical analysts on 20 questions: 90.0% EX, 90.0% JAC, 95.0% RQR. The sample is small; the 95% CI on EX is about [68%, 99%], so it is a directional reference. The quiz did not allow abstention, and both analysts produced executable SQL, so expert SR is undefined and SER is 0.
Under the baseline single-turn prompt, Gemini-3-Pro leads at 58.1% EX, 62.4% JAC, 81.8% RQR, and 0% syntax errors. Claude-4.5-Opus scores 54.8% EX and 80.6% RQR; GPT-o3-mini 53.5%; GPT-5.2 only 48.5%. Among open models, Qwen-2.5-Coder-32B reaches 40.8% EX, ahead of Llama-3.1-405B at 38.1%, but a 15.7% SER inflates its 61.0% Safety Rate: some abstentions are failed queries.
| System | EX | JAC | RQR |
| Domain expert (20 items) | 90.0% | 90.0% | 95.0% |
| Gemini-3-Pro baseline | 58.1% | 62.4% | 81.8% |
| BMSQL-GPT-o3-mini | 62.6% | 69.2% | 83.2% |
| DAIL-SQL-GPT-o3-mini | 61.2% | 63.6% | 81.4% |
| GPT-o3-mini baseline | 53.5% | 60.4% | 73.3% |
| Qwen-2.5-Coder-32B | 40.8% | 44.4% | 58.2% |
Few-shot helps. Ten examples lift GPT-o3-mini by 7.8 EX points to 61.3%. Forty shots add almost nothing. Sample rows alone barely move the needle; the bottleneck is schema, not memorizing cell values.
Interaction results split. Schema indexing is the weak setup: Index-GPT-4o scores 25.5% EX with 27.5% SER. ReAct helps GPT a little (GPT-o3-mini to 56.2% EX) and does not transfer cleanly across models. BMSQL-GPT-o3-mini records the best EX at 62.6%, statistically tied with DAIL-SQL-GPT-o3-mini at 61.2%. Paired with Gemini, BMSQL drops to 55.9% EX, below the Gemini-3-Pro single-turn baseline. BMSQL-Gemini reaches 84.6% RQR, still about ten points under the expert 95%. One BMSQL-GPT-o3-mini call uses about 39k tokens; Gemini-3-Pro baseline uses about 3,100.
The two systems fail in different places. ReAct rarely picks the wrong table and produces no syntax errors. BMSQL is better at writing p-value cutoffs and trial-phase filters into WHERE. Wrong table is the most common error overall, then missing or incorrect thresholds. Join, similarity search, and multi-filter queries are the hard bins.
Extra inference passes do little. From 1-pass to 3-pass, BMSQL-GPT-o3-mini EX goes from 62.6% to 61.7% while RQR goes from 83.2% to 85.5%; the extra tokens mostly fix syntax. On a 20-table schema, 10-shot falls from 61.3% to 53.8% EX (down 7.5 points); BMSQL falls from 62.6% to 58.4% (down 4.2). On the original schema, 10-shot already hits 61.3%, nearly matching BMSQL at about one-seventh the tokens.
The expert gap is about 25 to 30 EX points.
For anyone shipping text-to-SQL or science agents, this is a messier testbed than Spider or BIRD: BigQuery dialect, implicit statistical conventions, and a second score on whether the model can narrate the result. 58% execution accuracy does not belong in a discovery pipeline. The ethics note is blunt: a plausible query can still return the wrong set, and anything near clinical use needs a human check.
BMSQL is not a required architecture. The paper never evaluated open-ended coding agents such as Claude Code or Gemini CLI, and says the numbers should not be read as proof that a custom multi-step stack is necessary. Ten-shot already hits 61.3% on the 10-table schema at far lower cost. The two hard skills are separate: picking the right table, and encoding domain statistical conventions as filters.
Sixty-eight thousand items from 40 templates are more homogeneous than real questions. The authors cite GSM-Symbolic and long-tail entity effects to argue that swapping gene and disease names still stresses models. The 546-item test set is still sampled from those 40 templates, so the generalization radius is those 40 question families. Gold SQL is not unique, which is why JAC and RQR sit beside EX.
The expert baseline is 20 questions. The CI is wide. No inter-annotator agreement was computed. BigQuery locks out many general text-to-SQL toolchains; a SQLite port is planned. Open-ended coding agents were not run. The 20-table robustness gap sits inside the confidence intervals; the paper treats it as a trend, not a statistical claim.