Virtual Cell Benchmarks May Be Flawed
BoWang87 · x · 2026-07-08
A paper targeting ICML 2026 points out that standard perturbation benchmarks for virtual cell tasks might have severe issues.
The authors found that simple mean predictors outperform many models on these benchmarks. They argue the issue isn't poor model performance, but rather the evaluation method treating numerous irrelevant genes equally to a few signal genes, leading to distorted scores.
More from Research
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- Apodex Launches TRACES, First Benchmark for Evaluating 'Discoverative AI' on Real-World Problems — Faheem_uh · 2026-09-11
- TRACES grades the process, not the answer: six-dimension eval for open-ended AI science — Faheem_uh · 2026-09-11
- Apodex launches TRACES, a benchmark grading AI on open-ended discovery instead of known answers — Faheem_uh · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11