CMU's WhatWorkedBench Measures How Well AI Research Agents Understand Their Experiments

CarnegieMellonU · hf · 2026-09-24

CMU researchers introduce WhatWorkedBench, a benchmark for "experimental understanding": an agent's ability to predict how component changes affect outcomes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface predicting scores for every configuration; exhaustive CPU execution supplies reference effects.

Original post →

More from coding & agent

coding & agent channel →