Stanford's Kundaje Attacks Single-Cell Foundation Models and 'Virtual Cell' Hype in 12-Tweet Thread
On October 6, Stanford professor Anshul Kundaje posted a 12-tweet thread systematically critiquing a new benchmark study on perturbation prediction with single-cell foundation models. His core conclusion: current scFM approaches fall far short of delivering on the "virtual cell" promise.
Confirmed
- Kundaje noted that the only result the new paper demonstrates is that some scFMs explicitly fine-tuned on perturbation data can beat "utterly trivial" baselines under updated evaluation metrics—far too weak, in his view, to support the claim that scFMs are a path toward perturbation prediction or causal transcriptional regulation models.
- His central methodological argument: scFMs pretrained on observational data cannot, and will not, learn causal effects in a reliable way, and thus cannot be directly applied to perturbation prediction or causal discovery.
- He described the absolute performance of fine-tuned scFMs on perturbation data as "abysmal."
- He pointed out the paper omits some simple tweaks to trivial baselines (e.g., weighted means, or linear models trained directly against WMSE) that could strengthen them further, eroding the relative advantage of fine-tuned scFMs.
- He enumerated the perturbation prediction models that currently perform "reasonably well"—PRESAGE, Rheister, Arc's PIE model, GenbioAI's nearest-neighbor model—none of which use single-cell pretraining.
- SOTA perturbation models appear to predict unseen perturbations, but generalization largely holds only for perturbations and conditions highly similar to the training distribution; beyond the coverage of training data and prior knowledge, generalization remains poor. He argued this is not itself the problem, but users must understand this limitation.
- He went further: the bigger problem is data design—with current experimental designs and measurable inputs, it is fundamentally impossible to learn a "virtual cell," a world model, or a model with biological causal mechanisms.
- Since most of the predictive power of SOTA perturbation models comes from prior-knowledge information, he argued, calling them causal transcriptional regulation models or "virtual cells" is unfounded.
Unconfirmed
- Bo Wang's team published a companion paper in Nature Biotechnology, "Deep Learning Perturbation Models Can Outperform Baselines on Calibrated Metrics," arguing that "single-cell models can't beat baselines" is a measurement problem and that models can win under calibrated metrics—effectively a rebuttal. Kundaje's repost also suggested the dispute may come down to metrics. Which side is right remains to be settled by the community.
Why it matters
- The debate strikes at the scientific foundations of the hot "virtual cell" narrative in AI for Science: if foundation models trained on observational data cannot learn causal effects, funding and messaging around this direction may need rethinking.
- In his closing posts (posts 12 and 10–11), Kundaje criticized those who ignore empirical lessons and twist rigorous research into support for their own agenda, highlighting tension between evaluation rigor and science communication.
- Bo Wang's team's metrics-calibration paper in a Nature journal gives the criticized side a path to vindication; the clash will shape how future benchmarks for perturbation prediction models are designed.
2026-10-06 ~ 2026-10-06 · 12 related posts
Primary sources
- Kundaje: scFMs trained on observational data cannot learn causal effects reliably — anshulkundaje ·
- Kundaje: best-performing perturb models (PRESAGE, PIE, Rheister) use no scFM pretraining — anshulkundaje ·
- Nature Biotech paper: "single-cell models don't beat baselines" was a measurement problem — BoWang87 ·
- [source] Nature Biotech paper: "single-cell models don't beat baselines" was a measurement problem — BoWang87 · 2026-10-06
- Virtual cell debate was a metrics problem: Nature Biotech paper says calibrated metrics vindicate perturbation models — anshulkundaje · 2026-10-06
- Kundaje mocks benchmark: fine-tuned scFMs only beat the most trivial baselines — anshulkundaje · 2026-10-06
- [source] Kundaje: scFMs trained on observational data cannot learn causal effects reliably — anshulkundaje · 2026-10-06
- Kundaje: unreported baseline tweaks like weighted means could beat fine-tuned scFMs — anshulkundaje · 2026-10-06
- [source] Kundaje: best-performing perturb models (PRESAGE, PIE, Rheister) use no scFM pretraining — anshulkundaje · 2026-10-06
- Stanford's Kundaje: SOTA perturb models' power comes from prior knowledge, not virtual cells — anshulkundaje · 2026-10-06
- SOTA perturb models struggle to generalize beyond training perturbations, says Kundaje — anshulkundaje · 2026-10-06
- Stanford's Kundaje: perturbation models barely generalize beyond training distribution — anshulkundaje · 2026-10-06
- Kundaje: current experiment designs can't yield virtual cells or causal world models — anshulkundaje · 2026-10-06
- Kundaje wraps thread: folks spin rigorous study back into 'virtual cell' propaganda — anshulkundaje · 2026-10-06
1 near-duplicate retellings: anshulkundaje