Frontier models' research taste doubles every ~3 months, Opus 5.5 beats human experts 2.3x
ZeroStateReflex · x · 2026-10-07
pzeroresearch quantifies AI's 'experimental research taste': since December 2025, frontier models' ability has doubled roughly every 3 months, and the best model, Opus 5.5, now exceeds their expert human baseline. Opus 5.5 shows 2.3x the research taste of their best human experts — matching an expert's score with 17 GPU hours of experiments instead of 40.
Why it matters: in their AI Futures Model, research taste largely determines how quickly superintelligence arrives once coding is fully automated. Caveat: their human experts are experienced researchers, but most haven't worked at a frontier lab.
More from Models
- d1-3B runs on a MacBook: hands-on demos show Liquid AI's open model in action — JosephJacks_ · 2026-10-08
- SentenceTransformers gets native ColPali model support thanks to tomaarsen — tomaarsen · 2026-10-08
- ChatGPT Plus users report 'thinking' time doubled in 2025 with no quality gain — Gazialp · 2026-10-08
- Mistral Large 4 debuts at #45 on Code Arena WebDev, near Opus 4.8 at 1/6 the price — arena · 2026-10-08
- Liquid AI releases Open d1: open-weight 3B and 600M multimodal decision models — JosephJacks_ · 2026-10-08
- Liquid AI details d1-omni-600M: 600M params for text+image or text+audio — JosephJacks_ · 2026-10-08