Anthropic's internal model scores leak: new model may surpass Mythos 5
scaling01 · x · 2026-08-15
According to Anthropic's internal AECI evaluation, its new model (Model 2) scores 1.5 points higher than Mythos 5, estimated at 162.79. Based on the trendline, Anthropic may have reached that score around July 8, with training likely completed in May or June. Current internal models may score between 163-165, with latest untested models around 164.5-167.5.
More from Models
- Grok 4.6 now available in GitHub Copilot — intellectronica · 2026-08-15
- Gemini 3.7 Flash wins THOR Finding Triage Benchmark — zacharynado · 2026-08-15
- AI lab leaders don't worry about context windows; one thread hits billions of tokens — alliekmiller · 2026-08-15
- Gemini touted as the best model for Chess — Last_Conclusion_8984 · 2026-08-15
- GPT-5.6-sol leaks internal chain-of-thought during command execution — moyix · 2026-08-15
- Running Qwen3.8-27B on 2x3090: 200K Context with F16 KV, Vision, and Reasoning — Sisuuu · 2026-08-15