Calibrate multiomics with reference materials before feeding AI: a common ruler makes batches comparable

2026-09-03

An 18-institution team led by Fudan University proposes in Nature Biotechnology that standard reference materials be co-profiled with study samples, with results reported as sample-to-reference ratios (SRR), so multiomics data become reproducible across batches, labs and platforms and fit for AI training.

What problem this solves

AI is eating multiomics data. A single transcriptomic, proteomic or metabolomic measurement yields tens of thousands of features per sample, exactly the high-dimensional input models want. But the data has a metrology blind spot: the same biological sample produces different raw signals (intensity, I) on two instruments, in two labs or across two batches, in unpredictable ways. Training on such data teaches the model to learn instrument fingerprints as if they were biology.

This is not news. The MAQC (MicroArray Quality Control) consortium, founded at the US FDA in 2006, has repeatedly shown batch effects to be the leading root cause of irreproducibility in omics. The dominant fix has been computational correction after the fact (ComBat, limma, RUV), each of which assumes batch structure can be estimated from within the data itself. A a systematic comparison (Molania et al., Nat. Biotechnol. 41, 82–95, 2023) showed that without an external anchor these methods often remove real biological signal along with the batch. The "for artificial intelligence" in the title makes the stakes explicit: the stronger the downstream model, the more measurement error upstream gets amplified into systematic bias.

Method

The core mechanism in one sentence: give every batch a common ruler, report results as how many times the ruler the sample is, and the instrument sensitivity factor cancels in the division.

In analytical chemistry this is old discipline: numbers from an uncalibrated instrument do not count. In clinical testing, the ICH M10 guideline (the international harmonized standard for bioanalytical method validation) already mandates internal-standard ratio reporting. The paper's move is to bring that discipline into multiomics and argue why downstream AI needs it.

Results

This is a four-page comment (Brief Communication): a proposal plus a conceptual Fig. 1, no new experiments, no baseline table. What it offers instead is a chain of argument assembled from two decades of evidence:

ClaimEvidence cited
Batch effects are the root cause of omics irreproducibilityMAQC series (Shi et al., BMC Bioinformatics 6:S11, 2005; Nat. Biotechnol. 35:1127, 2017)
Purely computational correction fails without an external anchorMolania et al., Nat. Biotechnol. 41:82–95, 2023
The metrology approach has mature precedentICH M10 bioanalytical guideline (2023)
Chronic-disease precision medicine needs comparable dataZheng et al., Nat. Biotechnol. 42:1133–1149, 2024

The paper gives no head-to-head comparison of SRR against ComBat-style correction, and no quantified cross-lab SRR comparability numbers. Those are said to live in ref. 10 (Liu et al., bioRxiv 2026.08.09.740864); the preprint was blocked by Cloudflare and the comment sits behind a paywall, so the details could not be verified.

Why it matters

For anyone doing AI4Science, this is a data-supply-side comment. Three points land directly:

Verdict: a concept-proposal comment, not a methods paper. The concept, calibrate at measurement rather than correct downstream, holds up, backed by twenty years of MAQC/Quartet data. For AI practitioners it reads as a reminder to check whether training data were calibrated, not as a drop-in tool.

Limitations

Terms

Source

What people are saying

All paper explainers