Naive baseline beats AI model in ~95% of cases, exposing flaws in biology benchmarks
bravo_abad · x · 2026-10-07
Miller and colleagues show that AI benchmarks in cell biology can miss the very signal they're meant to detect.
- Setup: cells receiving the same genetic intervention were split into two groups; one group's average gene-expression profile was used to predict the other.
- Finding: in one dataset, a generic average that ignored the intervention entirely scored better in 95% of cases under mean-squared error.
- Why: most genes barely respond to interventions, so pooling many cells yields a cleaner estimate of the unchanged background; when errors are averaged across genes, this low-noise advantage outweighs capturing the actual response.
- Fix: the researchers built a positive control combining observed responses with the stable average where evidence of change was weak, enabling validation of scoring rules, and propose weighted MSE that emphasizes affected genes.
Takeaway: evaluation metrics for AI in biology may systematically reward predicting the mean and hide real intervention effects.
More from Research
- Gary Marcus: Vague AI Math Proof Report Would Never Pass Peer Review — GaryMarcus · 2026-10-08
- Dev Says He Found a Hidden Exoplanet in NASA Data Using Claude Code — iannuttall · 2026-10-08
- NeurIPS Paper LADE Detects Harmful Queries from First-Token Probabilities — mohitban47 · 2026-10-08
- Researchers Use AI to Design mRNA Vaccines Stable at Room Temperature for Up to a Year — Polymarket · 2026-10-08
- MIT's SOLE-R1: Video-Language Reasoning as the Sole Reward Enables Zero-Shot On-Robot RL — micoolcho · 2026-10-08
- Microsoft open-sources Agent Lightning v1.0: 3,500-line RL framework boosting SWE-bench +14.6 pts — Microsoft Research · 2026-10-08