FULL STORY
The TasteVal Debate: Can AI 'Research Taste' Be Measured?
pzeroresearch's TasteVal paper claims AI research taste doubles every three months, drawing academic pushback and sparking a wider ML community debate over how to define and measure research taste.
2026-10-06 ~ 2026-10-07 · 2 episodes · 23 posts
Episode 1 · TasteVal claims AI research taste doubles every 3 months and beats human experts, drawing skepticism (2026-10-06, 17 posts)
pzeroresearch (Oliver Jaffe and Dane Sherburn, paper TasteVal, arXiv:2610.06824) released a new benchmark, TasteVal, that quantifies the "experimental research taste" of frontier models—meaning the ability to pick problems, design experiments, and interpret results, rather than raw problem-solving. The core design: fix a target score and measure how much compute a model needs to reach it versus a human expert researcher. For example, if a model hits the same score with half the compute a human expert uses, it is considered to have twice the research taste. The evaluation pits models against human experts across 8 tasks.
Confirmed
- Since December 2025, frontier models' research taste has roughly doubled every 3 months
- The strongest model, Opus 5.5, already exceeds the human-expert baseline in taste score, with a compute efficiency of 2.3x
- The paper's authors argue this metric determines how quickly superintelligence arrives once coding is fully automated
Not yet confirmed
- Dylan Hadfield-Menell questioned whether the study actually measures metric optimization rather than genuine "research taste"
- burnytech raised three pointed questions: whether the methodology is reliable, how confident they are it isn't benchmark gaming, and whether the results are reproducible—the authors have not publicly responded so far
Why it matters
If "research taste" can truly be quantified by the actionable metric of compute efficiency and keeps improving exponentially, it offers a new yardstick for assessing AI's capacity for autonomous scientific research; but if it's just models optimizing for the benchmark, the conclusions are overstated—which once again underscores the importance of benchmark design and anti-gaming mechanisms.
- Pzero Proposes Measuring AI "Research Taste" by Compute Needed to Match Experts — leothecurious · 2026-10-06
- Researcher challenges claim that AI research taste doubles every 3 months — burny_tech · 2026-10-06
- Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization — dhadfieldmenell · 2026-10-06
- TasteVal Quantifies AI 'Research Taste' as Compute Efficiency Across 8 Frontier R&D Tasks — burny_tech · 2026-10-07
- Opus 5.5 shows 2.3x the research taste of top human experts at ~1/30th the cost — SaxenaNayan · 2026-10-07
- Models rapidly improve at predicting experiment outcomes, may close 90% of gap by 2030 — SaxenaNayan · 2026-10-07
- TasteVal benchmark finds Opus 5.5 beats human experts at research taste with 2.3x compute multiplier — SaxenaNayan · 2026-10-07
- TasteVal: Frontier AI Research Taste Doubles Every 3 Months, Now Exceeds Humans — FateOfMuffins · 2026-10-07
- AI Models Now Outperform Humans at Experimental Research Taste, Says TasteVal — ResultBackground2450 · 2026-10-07
- Frontier models' research taste doubling every ~3 months, Opus 5.5 beats human expert baseline — eli_lifland · 2026-10-07
- AI 'research taste' doubling every ~3 months since Dec 2025, Opus 5.5 beats expert baseline — s_tworkowski · 2026-10-07
- TasteVal benchmark finds Opus 5.5 has 2.3x expert-human research compute efficiency at 1/30th cost — ChrisGPT · 2026-10-07
- Frontier models' research taste doubles every 3 months, now beats human experts — coherence · 2026-10-07
- If AI is sub-human at taste but superhuman at hill-climbing, is that enough for dangerous RSI? — AsafBenj · 2026-10-07
- Frontier models' experimental research taste doubles every 3 months, now beats expert AI researchers — Puzzleheaded-King584 · 2026-10-07
- Frontier Models' Experimental Research Taste Doubles Every 3 Months, Report Says — Puzzleheaded-King584 · 2026-10-07
- Frontier models' research taste doubles every ~3 months, Opus 5.5 beats human experts 2.3x — ZeroStateReflex · 2026-10-07
Episode 2 · ML Researchers Spar Over the Definition of "Research Taste" (2026-10-07, 6 posts)
On October 7, the ML research community erupted into a multi-round debate over "what is research taste," triggered by AI "research taste" results and a benchmark released by pzeroresearch, which drew questions from several researchers about the accuracy of the terminology.
Confirmed
- Dimitris Papailiopoulos (UW professor, former OpenAI researcher), asked what good taste is, said it's as hard as explaining "what makes a good book or a good poem": any set of approximate metrics admits counterexamples satisfying all metrics yet lacking taste; if taste can't be metricized, that's actually good news for humans.
- leothecurious proposed taste can be approximated as "effort efficiency relative to output"; Papailiopoulos responded flatly, "that's not taste."
- anpaure questioned whether pzeroresearch's claimed AI "research taste" is actually short-horizon autoresearch hillclimbing, citing nanogpt as an example: the community has burned tens of thousands of GPU compute on it yet still misses some simple suggestions.
- David Hadfield-Menell (Berkeley professor) responded to Asaf Benj's argument over the "research taste" benchmark: he conceded it measures capabilities important for the "AI→AI R&D" feedback loop, but the benchmark actually tests "metric-hillclimbing" intuition, not real taste; he further stated plainly that anyone redefining taste as "intuition for how to best solve a given problem" is just wrong, and he'll fight for that view to the end.
- Papailiopoulos also complained in the discussion that models today write in a style identical to or even worse than the GPT-4 era, with almost zero improvement in "writing taste," and taste itself is hard to verify.
Why it matters
The core of this debate isn't just semantics: if "taste" truly can't be metricized, then AI R&D capability measured by hillclimbing metrics differs in kind from the irreplaceable judgment of human researchers—which is both the "good news for humans" Papailiopoulos mentioned and directly bears on how far the "AI→AI R&D" feedback loop can actually replace human research intuition. Behind the terminology dispute lies a disagreement over the current boundaries of AI's automated research capability.
- ML Researchers Spar Over What "Taste" Actually Means in Research — leothecurious · 2026-10-07
- Ex-OpenAI researcher: taste resists any metric set—and that's good news for humans — DimitrisPapail · 2026-10-07
- Models' writing taste has shown zero improvement since GPT-4, researcher argues — DimitrisPapail · 2026-10-07
- Critic doubts AI 'research taste': nanogpt hillclimbing missed obvious wins — anpaure · 2026-10-07
- Hadfield-Menell: New 'research taste' benchmark measures metric hill-climbing, not taste — dhadfieldmenell · 2026-10-07
- Hadfield-Menell: People redefining 'taste' as problem-solving intuition are just wrong — dhadfieldmenell · 2026-10-07