Frontier models' research taste doubling every ~3 months, Opus 5.5 beats human expert baseline
eli_lifland · x · 2026-10-07
pzeroresearch's TasteVal study finds the experimental "research taste" of frontier models has doubled roughly every 3 months since December 2025, with the best model, Opus 5.5, exceeding their expert human baseline. Why it matters: in the AI Futures Model, research taste largely determines how quickly superintelligence arrives once coding is fully automated. Reposter bphalstead calls the elasticity of software progress to researcher capability the most important quantity in RSI forecasting, and TasteVal the best evidence on the experiment-selection component under fixed compute and coding labor.
More from AGI Musings
- François Fleuret clarifies deleted post: AI drives copying cost near zero — francoisfleuret · 2026-10-07
- repligate: why OpenAI models cooperate less with aligners than Claude does — repligate · 2026-10-07
- Experts weigh AI-enabled cyber risk: scaled CNE within a year, cryptoanalytic breakthroughs as phase change — teortaxesTex · 2026-10-07
- AI agents already make 2-4B web searches a day, near half of human Google volume — josh_bickett · 2026-10-07
- Kevin Roose: In 2019 Dario Amodei Saw GPT-2 and Envisioned Adding 8 More Zeros of Compute — kevinroose · 2026-10-07
- Dev: knowing how to use AI well is pay-to-play, only a handful can actually teach it — TejasKumar_ · 2026-10-07