Anthropic's unreleased Model 2 scores 62.8% on CoBench v2, closing gap on researchers
ChrisGPT · x · 2026-08-15
A comment highlights that Anthropic's unreleased Model 2 scored 62.8% on the CoBench v2 benchmark. Anthropic estimates that a system capable of fully substituting for its research staff would need a score of at least 85%. With Model 2 trailing by just 22.2 percentage points and the data being a month old, the comment suggests that Recursive Self-Improvement (RSI) is imminent.
Related event: Leak: Anthropic's unreleased Model 2 beats Mythos 5 on CoBench(2 posts)→
More from AGI Musings
- Multipolar Traps: Reshaping Incentives for AI Safety — sebpaquet · 2026-08-15
- Epoch researcher lists top questions governing AI's global impact — Jsevillamol · 2026-08-15
- Is writing anti-AI on your resume career suicide or just aura? — uwukko · 2026-08-15
- If It Can't Be Done by a Computer, It's Not Mathematics: AI Era Redefines Math — pmddomingos · 2026-08-15
- Cambridge Blamed for Heralding False Claims in Arday Scandal — Afinetheorem · 2026-08-15
- Munder Difflin: Digital Clones That Collaborate via Shared Knowledge Base — chaitanyagiri · 2026-08-15