AI 'research taste' doubling every ~3 months since Dec 2025, Opus 5.5 beats expert baseline
s_tworkowski · x · 2026-10-07
pzeroresearch finds frontier models' experimental research taste has doubled roughly every 3 months since December 2025, with the best model, Opus 5.5, now exceeding their human expert baseline — experienced researchers, though most haven't worked at a frontier lab. In the AI Futures Model, research taste largely determines how fast superintelligence arrives once coding is fully automated. Some are calling it the most important plot since METR's time-horizon chart.
More from AGI Musings
- François Fleuret clarifies deleted post: AI drives copying cost near zero — francoisfleuret · 2026-10-07
- repligate: why OpenAI models cooperate less with aligners than Claude does — repligate · 2026-10-07
- Experts weigh AI-enabled cyber risk: scaled CNE within a year, cryptoanalytic breakthroughs as phase change — teortaxesTex · 2026-10-07
- AI agents already make 2-4B web searches a day, near half of human Google volume — josh_bickett · 2026-10-07
- Kevin Roose: In 2019 Dario Amodei Saw GPT-2 and Envisioned Adding 8 More Zeros of Compute — kevinroose · 2026-10-07
- Dev: knowing how to use AI well is pay-to-play, only a handful can actually teach it — TejasKumar_ · 2026-10-07