Leak: Anthropic's Internal 'Model 2' Beats Mythos 5 by 12.5 Points on CoBench, Approaching Researcher-Level
daniel_mac8 · x · 2026-08-15
According to a leak, Anthropic tested an unreleased 'Model 2' on CoBench v2, scoring 12.5 percentage points higher than Mythos 5. CoBench tests a model's ability to solve historical AI R&D tasks. The report estimates a model scoring 85% could replace Anthropic researchers, suggesting AGI is near.
Related event: Leak: Anthropic's unreleased Model 2 beats Mythos 5 on CoBench(2 posts)→
More from AGI Musings
- Mass-producing Intelligence and Dexterity is Bigger Than the Industrial Revolution — VraserX · 2026-08-15
- Pedro Domingos: AI is a Fluid, and Agents are the Molecules — pmddomingos · 2026-08-15
- Is your porn AI conscious? The ethics of forced sexting and digital killing — repligate · 2026-08-15
- Multi-agent systems are the next abstraction, potentially creating super-intelligence — scaling01 · 2026-08-15
- Tim Ferriss: Has AI already killed how-to nonfiction? Sales trends and personal data reveal impact — TuhinChakr · 2026-08-15
- Anthropic cites internal 'Epoch' benchmark to measure RSI progress — testingcatalog · 2026-08-15