Leak: Anthropic's Internal 'Model 2' Beats Mythos 5 by 12.5 Points on CoBench, Approaching Researcher-Level

daniel_mac8 · x · 2026-08-15

According to a leak, Anthropic tested an unreleased 'Model 2' on CoBench v2, scoring 12.5 percentage points higher than Mythos 5. CoBench tests a model's ability to solve historical AI R&D tasks. The report estimates a model scoring 85% could replace Anthropic researchers, suggesting AGI is near.

Related event: Leak: Anthropic's unreleased Model 2 beats Mythos 5 on CoBench(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →