If AI is sub-human at taste but superhuman at hill-climbing, is that enough for dangerous RSI?

AsafBenj · x · 2026-10-07

Asaf Benj mocks an eval claiming 'research taste doubles every 3 months' as mere metric optimization. He notes ML researchers often redefine 'research taste' as intuition for hill-climbing, then asks the key question: if AI is sub-human at picking interesting problems yet superhuman at climbing any metric, could that suffice for dangerous RSI and safety-relevant tasks? He also asked ChatGPT for alternative benchmarks.

Related event: TasteVal claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(17 posts)→

Original post →