Frontier models' research taste doubles every ~3 months, Opus 5.5 beats human experts 2.3x

ZeroStateReflex · x · 2026-10-07

pzeroresearch quantifies AI's 'experimental research taste': since December 2025, frontier models' ability has doubled roughly every 3 months, and the best model, Opus 5.5, now exceeds their expert human baseline. Opus 5.5 shows 2.3x the research taste of their best human experts — matching an expert's score with 17 GPU hours of experiments instead of 40.

Why it matters: in their AI Futures Model, research taste largely determines how quickly superintelligence arrives once coding is fully automated. Caveat: their human experts are experienced researchers, but most haven't worked at a frontier lab.

Related event: TasteVal claims AI research taste doubles every 3 months and beats human experts, drawing skepticism(17 posts)→

Original post →

More from Models

Models channel →