Miscalibrated metrics skewed benchmarks: deep learning gene perturbation models do beat baselines
simocristea · x · 2026-10-02
A newly peer-reviewed Nature Biotechnology brief challenges prior benchmarks claiming deep-learning genetic perturbation models fail to beat uninformative baselines. The authors introduce a positive-control baseline and metric calibration framework, showing across 14 datasets and 18 metrics that common metrics like MSE and Pearson Δ are frequently miscalibrated and insensitive to true model performance. Under well-calibrated metrics, deep-learning perturbation models can indeed outperform uninformative baselines.
Related event: Nature Biotechnology paper challenges gene perturbation model benchmarks(2 posts)→
More from Research
- KL divergence between Cauchy distributions is finite and symmetric, with closed form — FrnkNlsn · 2026-10-02
- Learner to share daily notes through Stanford's CS224n NLP course — stanfordnlp · 2026-10-02
- Building CS224n's Transformer from scratch: 144 lines expose scaling and mask bugs — stanfordnlp · 2026-10-02
- One human health check erases semaglutide's behavioral effect in mice, preprint finds — BenSiranosian · 2026-10-02
- BeyondSCe grasps objects by past-event reference, hitting 77% success on occluded targets — SeoulNationalUniv · 2026-10-02
- RouteFM pretrains LLM routing once for anywhere: +2.23 quality points over strongest baseline — nanjinguniv · 2026-10-02