Frontier models' research taste doubling every ~3 months, Opus 5.5 beats human expert baseline

eli_lifland · x · 2026-10-07

pzeroresearch's TasteVal study finds the experimental "research taste" of frontier models has doubled roughly every 3 months since December 2025, with the best model, Opus 5.5, exceeding their expert human baseline. Why it matters: in the AI Futures Model, research taste largely determines how quickly superintelligence arrives once coding is fully automated. Reposter bphalstead calls the elasticity of software progress to researcher capability the most important quantity in RSI forecasting, and TasteVal the best evidence on the experiment-selection component under fixed compute and coding labor.

Related event: TasteVal benchmark claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(12 posts)→

Original post →

More from AGI Musings

AGI Musings channel →