Weighted eval metrics for gene-effect models spark debate over circular benchmark design
anshulkundaje · x · 2026-10-06
Genomics ML researcher Anshul Kundaje traded technical points with a critic over new evaluation metrics for gene differential-effect models:
- The critic argued the new weighted metrics depend on the very measurements they evaluate, a circular scoring concern
- Kundaje replied that the weights penalize errors more on genes with stronger differential-effect evidence and are derived from the data itself, which he considers acceptable
- He added models should ideally train with losses that upweight signals they care about — meaning baselines could arguably improve simply by training with a WMSE loss
A substantive exchange on benchmark methodology: metric design and training objectives that don't align can distort cross-model comparisons.
More from Research
- Optimized open-source Argus hits SOTA on WGO-Bench, lifting semantic F1 from 29.8% to 50.2% — _sonith · 2026-10-06
- Alison Gopnik's 'Explanation as Orgasm' Hypothesis Goes Viral — stevenstrogatz · 2026-10-06
- First NeurIPS 2026 Agent Behavior Workshop Accepts 149 Papers — DanielKhashabi · 2026-10-06
- Black Forest Labs open-sources 7B FLUX 3 Action, tops RoboLab-120 at 3.95x speed — dl_weekly · 2026-10-06
- SwiLA explained: a mixture of J linear regressions that reduces to DeltaNet at J=1 — YouJiacheng · 2026-10-06
- MIT paper: stating rules beats reasoning prompts for LLM agents — Gemma's underbid gap shrinks from $5.30 to $0.30 — rohanpaul_ai · 2026-10-06