Anthropic's unreleased Model 2 scores 62.8% on CoBench v2, closing gap on researchers

ChrisGPT · x · 2026-08-15

A comment highlights that Anthropic's unreleased Model 2 scored 62.8% on the CoBench v2 benchmark. Anthropic estimates that a system capable of fully substituting for its research staff would need a score of at least 85%. With Model 2 trailing by just 22.2 percentage points and the data being a month old, the comment suggests that Recursive Self-Improvement (RSI) is imminent.

Related event: Leak: Anthropic's unreleased Model 2 beats Mythos 5 on CoBench(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →