Anthropic Report Reveals Internal Model Exceeding Mythos 5
In its second 186-page Risk Report, Anthropic disclosed that a model with the internal codename "Model 2" slightly outperforms Claude Mythos 5 overall and is already heavily used for coding, data generation, research, and agentic tasks—though the company made clear there are "no plans to release it externally." The model scored 62.8% on the CoBench v2 test versus 50.3% for Mythos 5, a 12.5-percentage-point gap, but it still falls short of the 85% target for fully automated researcher replacement.
Confirmed
- Model 2 capabilities: Slightly better than Mythos 5 (Claude 3.5 Sonnet) overall, scoring 62.8% on internal code and infrastructure diagnostics tasks.
- Deployment status: Full pre-deployment evaluation not yet completed; currently limited to internal use.
- Performance comparison: On the CoBench v2 test, Model 2 outscored Mythos 5 by 12.5 percentage points.
Unconfirmed
- Architecture speculation: Some suggest Mythos Preview may be a roughly 10T-parameter "teacher" model, with Model 1 and Model 2 as iterative versions, and Mythos 5 and Fable 5 as distilled variants.
- Specific scores: Internal AECI evaluations show Model 2 scoring 1.5 points higher than Mythos 5, estimated at around 162.79 points.
2026-08-15 ~ 2026-08-15 · 11 related posts
Primary sources
- Anthropic confirms Mythos 5 as top internal model, hints at mysterious Model 2 — scaling01 · 2026-08-15
- Rumor: Anthropic uses internal model better than Mythos 5 — kimmonismus · 2026-08-15
- Rumor: Mythos Preview may be a 10T parameter teacher model — scaling01 · 2026-08-15
- [source] Anthropic Report Leaks 'Model 2', Slightly Stronger Than Claude 3.5 Sonnet — ChrisGPT · 2026-08-15
- Leak: Anthropic's Internal 'Model 2' Beats Mythos 5 by 12.5 Points on CoBench, Approaching Researcher-Level — daniel_mac8 · 2026-08-15
- Anthropic's internal model scores leak: new model may surpass Mythos 5 — scaling01 · 2026-08-15
- Anthropic's unreleased Model 2 scores 62.8% on CoBench v2, closing gap on researchers — ChrisGPT · 2026-08-15
- [source] Report: Anthropic's Internal 'Model 2' Outperforms Mythos 5 on Diagnostics — imjustnewatai · 2026-08-15
- Rumor: Anthropic Uses Internal Model Far Better Than Mythos 5 — Neurogence · 2026-08-15
- [source] Anthropic Reveals Unreleased Internal Model That Beats Mythos 5 in 186-Page Risk Report — 机器之心 · 2026-08-15
1 near-duplicate retellings: kimmonismus