TasteVal benchmark claims AI research taste doubles every 3 months and beats human experts, drawing skepticism
pzeroresearch (Oliver Jaffe and Dane Sherburn, paper TasteVal, arXiv:2610.06824) released a new benchmark, TasteVal, that quantifies the "experimental research taste" of frontier models—meaning the ability to pick problems, design experiments, and interpret results, rather than raw problem-solving. The core design: fix a target score and measure how much compute a model needs to reach it versus a human expert researcher. For example, if a model hits the same score with half the compute a human expert uses, it is considered to have twice the research taste. The evaluation pits models against human experts across 8 tasks.
Confirmed
- Since December 2025, frontier models' research taste has roughly doubled every 3 months
- The strongest model, Opus 5.5, already exceeds the human-expert baseline in taste score, with a compute efficiency of 2.3x
- The paper's authors argue this metric determines how quickly superintelligence arrives once coding is fully automated
Not yet confirmed
- Dylan Hadfield-Menell questioned whether the study actually measures metric optimization rather than genuine "research taste"
- burnytech raised three pointed questions: whether the methodology is reliable, how confident they are it isn't benchmark gaming, and whether the results are reproducible—the authors have not publicly responded so far
Why it matters
If "research taste" can truly be quantified by the actionable metric of compute efficiency and keeps improving exponentially, it offers a new yardstick for assessing AI's capacity for autonomous scientific research; but if it's just models optimizing for the benchmark, the conclusions are overstated—which once again underscores the importance of benchmark design and anti-gaming mechanisms.
2026-10-06 ~ 2026-10-07 · 12 related posts
Primary sources
- TasteVal: Frontier AI Research Taste Doubles Every 3 Months, Now Exceeds Humans — FateOfMuffins ·
- TasteVal benchmark finds Opus 5.5 beats human experts at research taste with 2.3x compute multiplier — SaxenaNayan ·
- Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization — dhadfieldmenell ·
- Pzero Proposes Measuring AI "Research Taste" by Compute Needed to Match Experts — leothecurious · 2026-10-06
- Researcher challenges claim that AI research taste doubles every 3 months — burny_tech · 2026-10-06
- [source] Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization — dhadfieldmenell · 2026-10-06
- TasteVal Quantifies AI 'Research Taste' as Compute Efficiency Across 8 Frontier R&D Tasks — burny_tech · 2026-10-07
- Opus 5.5 shows 2.3x the research taste of top human experts at ~1/30th the cost — SaxenaNayan · 2026-10-07
- Models rapidly improve at predicting experiment outcomes, may close 90% of gap by 2030 — SaxenaNayan · 2026-10-07
- [source] TasteVal benchmark finds Opus 5.5 beats human experts at research taste with 2.3x compute multiplier — SaxenaNayan · 2026-10-07
- [source] TasteVal: Frontier AI Research Taste Doubles Every 3 Months, Now Exceeds Humans — FateOfMuffins · 2026-10-07
- AI Models Now Outperform Humans at Experimental Research Taste, Says TasteVal — ResultBackground2450 · 2026-10-07
- TasteVal benchmark finds Opus 5.5 has 2.3x expert-human research compute efficiency at 1/30th cost — ChrisGPT · 2026-10-07
2 near-duplicate retellings: eli_lifland · s_tworkowski