ML Researchers Spar Over the Definition of "Research Taste"
On October 7, the ML research community erupted into a multi-round debate over "what is research taste," triggered by AI "research taste" results and a benchmark released by pzeroresearch, which drew questions from several researchers about the accuracy of the terminology.
Confirmed
- Dimitris Papailiopoulos (UW professor, former OpenAI researcher), asked what good taste is, said it's as hard as explaining "what makes a good book or a good poem": any set of approximate metrics admits counterexamples satisfying all metrics yet lacking taste; if taste can't be metricized, that's actually good news for humans.
- leothecurious proposed taste can be approximated as "effort efficiency relative to output"; Papailiopoulos responded flatly, "that's not taste."
- anpaure questioned whether pzeroresearch's claimed AI "research taste" is actually short-horizon autoresearch hillclimbing, citing nanogpt as an example: the community has burned tens of thousands of GPU compute on it yet still misses some simple suggestions.
- David Hadfield-Menell (Berkeley professor) responded to Asaf Benj's argument over the "research taste" benchmark: he conceded it measures capabilities important for the "AI→AI R&D" feedback loop, but the benchmark actually tests "metric-hillclimbing" intuition, not real taste; he further stated plainly that anyone redefining taste as "intuition for how to best solve a given problem" is just wrong, and he'll fight for that view to the end.
- Papailiopoulos also complained in the discussion that models today write in a style identical to or even worse than the GPT-4 era, with almost zero improvement in "writing taste," and taste itself is hard to verify.
Why it matters
The core of this debate isn't just semantics: if "taste" truly can't be metricized, then AI R&D capability measured by hillclimbing metrics differs in kind from the irreplaceable judgment of human researchers—which is both the "good news for humans" Papailiopoulos mentioned and directly bears on how far the "AI→AI R&D" feedback loop can actually replace human research intuition. Behind the terminology dispute lies a disagreement over the current boundaries of AI's automated research capability.
2026-10-07 ~ 2026-10-07 · 6 related posts
- Episode 1: TasteVal claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(2026-10-06, 17 posts)
- Episode 2: ML Researchers Spar Over the Definition of "Research Taste"(2026-10-07, 6 posts)
Primary sources
- ML Researchers Spar Over What "Taste" Actually Means in Research — leothecurious · 2026-10-07
- [source] Ex-OpenAI researcher: taste resists any metric set—and that's good news for humans — DimitrisPapail · 2026-10-07
- Models' writing taste has shown zero improvement since GPT-4, researcher argues — DimitrisPapail · 2026-10-07
- [source] Critic doubts AI 'research taste': nanogpt hillclimbing missed obvious wins — anpaure · 2026-10-07
- [source] Hadfield-Menell: New 'research taste' benchmark measures metric hill-climbing, not taste — dhadfieldmenell · 2026-10-07
- Hadfield-Menell: People redefining 'taste' as problem-solving intuition are just wrong — dhadfieldmenell · 2026-10-07