CMU's WhatWorkedBench Measures How Well AI Research Agents Understand Their Experiments
CarnegieMellonU · hf · 2026-09-24
CMU researchers introduce WhatWorkedBench, a benchmark for "experimental understanding": an agent's ability to predict how component changes affect outcomes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface predicting scores for every configuration; exhaustive CPU execution supplies reference effects.
- 36 tasks from 30 data sources and 8 workflow types, 1248 configuration records; core eval spans 4,206 numerical-control records and 108 agent episodes
- With 8 new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and caps effect error at 10% of score range on three
- Fitting a Gaussian process raises effect recovery from 0.632 to 0.698; encoding code equivalences lifts GP recovery from 0.248 to 0.462
- Designed for research on experimental agents, adaptive experimental design, and numerical inference
More from coding & agent
- A Hermes Agent prompt that audits Claude's memory files every session — alexcovo_eth · 2026-09-24
- Open-Source MCP Server Connects Claude/Cursor to ComfyUI with 50+ Tools — Less_Actuary_9441 · 2026-09-24
- Using Notion as a data hub makes switching AI services painless — ivanhzhao · 2026-09-24
- Hot take: a 22-year-old fluent in coding agents beats a lazy senior dev — but watch out for 'slop grenades' — jobergum · 2026-09-24
- Blender MCP hits 29k stars as Opus 5.5 builds cities on medium effort — sidahuj · 2026-09-24
- AI Agent Swarm Reverse-Engineers 2001 GBA Game Snood Byte-for-Byte in Two Weeks — Aizkmusic · 2026-09-24