Miscalibrated metrics skewed benchmarks: deep learning gene perturbation models do beat baselines

simocristea · x · 2026-10-02

A newly peer-reviewed Nature Biotechnology brief challenges prior benchmarks claiming deep-learning genetic perturbation models fail to beat uninformative baselines. The authors introduce a positive-control baseline and metric calibration framework, showing across 14 datasets and 18 metrics that common metrics like MSE and Pearson Δ are frequently miscalibrated and insensitive to true model performance. Under well-calibrated metrics, deep-learning perturbation models can indeed outperform uninformative baselines.

Related event: Nature Biotechnology paper challenges gene perturbation model benchmarks(2 posts)→

Original post →

More from Research

Research channel →