AI 'research taste' doubling every ~3 months since Dec 2025, Opus 5.5 beats expert baseline

s_tworkowski · x · 2026-10-07

pzeroresearch finds frontier models' experimental research taste has doubled roughly every 3 months since December 2025, with the best model, Opus 5.5, now exceeding their human expert baseline — experienced researchers, though most haven't worked at a frontier lab. In the AI Futures Model, research taste largely determines how fast superintelligence arrives once coding is fully automated. Some are calling it the most important plot since METR's time-horizon chart.

Related event: TasteVal benchmark claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(12 posts)→

Original post →

More from AGI Musings

AGI Musings channel →