Google's ScientistTwo solves 80.4% of 107 top-venue ML problems autonomously
thisdudelikesAI · x · 2026-09-22
Google unveiled ScientistTwo, an autonomous scientist pipeline: given a research problem, it surveys the state of the art, identifies limitations, generates hypotheses, writes code, runs experiments across datasets, performs ablations, drafts the paper, and iterates through a simulated peer-review loop where a rebuttal agent runs new experiments to answer reviewer criticisms. A meta-reviewer gates each round with no human in the loop after the problem statement.
Tested on 107 problems drawn from accepted ICLR, ICML, and NeurIPS papers:
- Advanced 86 of 107 (80.4% success rate)
- 25.2% average relative gain over human state-of-the-art
- Papers score above the average accepted ICLR 2026 and NeurIPS 2025 paper under the Stanford evaluation
More from AGI Musings
- "We've Reached Good Enough Intelligence" — Dev Says Focus on Speed, Not Smarter Models — charles_irl · 2026-09-22
- Token Demand to Bifurcate: "Good Enough" Models vs. Autonomous Research Labs — menhguin · 2026-09-22
- Critique: the closed circular world of AI Safety research, all funded by one guy — banteg · 2026-09-22
- If agents can shop for you, agent-native rivals will replace Zomato, Amazon and Blinkit — vaibhavbetter · 2026-09-22
- Book review: 'Silicon Sovereigns' examines AI, international law and the tech-industrial complex — ProfChesterman · 2026-09-22
- Academic argues substance and provenance matter more than whether text is AI-written — 3scorciav · 2026-09-22