If AI is sub-human at taste but superhuman at hill-climbing, is that enough for dangerous RSI?
AsafBenj · x · 2026-10-07
Asaf Benj mocks an eval claiming 'research taste doubles every 3 months' as mere metric optimization. He notes ML researchers often redefine 'research taste' as intuition for hill-climbing, then asks the key question: if AI is sub-human at picking interesting problems yet superhuman at climbing any metric, could that suffice for dangerous RSI and safety-relevant tasks? He also asked ChatGPT for alternative benchmarks.