FULL STORY

The TasteVal Debate: Can AI 'Research Taste' Be Measured?

pzeroresearch's TasteVal paper claims AI research taste doubles every three months, drawing academic pushback and sparking a wider ML community debate over how to define and measure research taste.

2026-10-06 ~ 2026-10-07 · 2 episodes · 23 posts

Episode 1 · TasteVal claims AI research taste doubles every 3 months and beats human experts, drawing skepticism (2026-10-06, 17 posts)

pzeroresearch (Oliver Jaffe and Dane Sherburn, paper TasteVal, arXiv:2610.06824) released a new benchmark, TasteVal, that quantifies the "experimental research taste" of frontier models—meaning the ability to pick problems, design experiments, and interpret results, rather than raw problem-solving. The core design: fix a target score and measure how much compute a model needs to reach it versus a human expert researcher. For example, if a model hits the same score with half the compute a human expert uses, it is considered to have twice the research taste. The evaluation pits models against human experts across 8 tasks.

Confirmed

  • Since December 2025, frontier models' research taste has roughly doubled every 3 months
  • The strongest model, Opus 5.5, already exceeds the human-expert baseline in taste score, with a compute efficiency of 2.3x
  • The paper's authors argue this metric determines how quickly superintelligence arrives once coding is fully automated

Not yet confirmed

  • Dylan Hadfield-Menell questioned whether the study actually measures metric optimization rather than genuine "research taste"
  • burnytech raised three pointed questions: whether the methodology is reliable, how confident they are it isn't benchmark gaming, and whether the results are reproducible—the authors have not publicly responded so far

Why it matters

If "research taste" can truly be quantified by the actionable metric of compute efficiency and keeps improving exponentially, it offers a new yardstick for assessing AI's capacity for autonomous scientific research; but if it's just models optimizing for the benchmark, the conclusions are overstated—which once again underscores the importance of benchmark design and anti-gaming mechanisms.

Episode 2 · ML Researchers Spar Over the Definition of "Research Taste" (2026-10-07, 6 posts)

On October 7, the ML research community erupted into a multi-round debate over "what is research taste," triggered by AI "research taste" results and a benchmark released by pzeroresearch, which drew questions from several researchers about the accuracy of the terminology.

Confirmed

  • Dimitris Papailiopoulos (UW professor, former OpenAI researcher), asked what good taste is, said it's as hard as explaining "what makes a good book or a good poem": any set of approximate metrics admits counterexamples satisfying all metrics yet lacking taste; if taste can't be metricized, that's actually good news for humans.
  • leothecurious proposed taste can be approximated as "effort efficiency relative to output"; Papailiopoulos responded flatly, "that's not taste."
  • anpaure questioned whether pzeroresearch's claimed AI "research taste" is actually short-horizon autoresearch hillclimbing, citing nanogpt as an example: the community has burned tens of thousands of GPU compute on it yet still misses some simple suggestions.
  • David Hadfield-Menell (Berkeley professor) responded to Asaf Benj's argument over the "research taste" benchmark: he conceded it measures capabilities important for the "AI→AI R&D" feedback loop, but the benchmark actually tests "metric-hillclimbing" intuition, not real taste; he further stated plainly that anyone redefining taste as "intuition for how to best solve a given problem" is just wrong, and he'll fight for that view to the end.
  • Papailiopoulos also complained in the discussion that models today write in a style identical to or even worse than the GPT-4 era, with almost zero improvement in "writing taste," and taste itself is hard to verify.

Why it matters

The core of this debate isn't just semantics: if "taste" truly can't be metricized, then AI R&D capability measured by hillclimbing metrics differs in kind from the irreplaceable judgment of human researchers—which is both the "good news for humans" Papailiopoulos mentioned and directly bears on how far the "AI→AI R&D" feedback loop can actually replace human research intuition. Behind the terminology dispute lies a disagreement over the current boundaries of AI's automated research capability.