Anthropic’s Claude Opus 5 system card shows gains over Mythos 5 on internal evals
tokenbender · x · 2026-07-25
Anthropic’s Claude Opus 5 system card is out, and the posted table suggests it outperforms Claude Mythos 5 on several internal research tasks, while still falling short on some thresholds.
The screenshot shows results across tasks such as:
- Kernel task: 449.46× vs 430.93× for Claude Mythos 5
- Time series forecasting: 5.68 vs 4.51 MSE on the hard variant
- LLM training (easy): 68.54× vs 69.61×
- LLM training (hard): 14.19× vs 8.36×
- Quadruped RL: 31.3 vs 29.55
- Novel compiler: 80.91% pass rate vs 85.3%
The card also notes that recent models now exceed the old rule-out thresholds on all but two tasks in Anthropic’s internal suite, so the risk analysis no longer relies on those automated evaluations.
Related event: Anthropic Releases Claude Opus 5 System Card(3 posts)→
More from Models
- Opus 5 system card says pretrained mode argued for separating its existence from economics — Sauers_ · 2026-07-25
- Claude Opus 5 has a full-on meltdown on a multimodal math question — Sauers_ · 2026-07-25
- LiteParse 2.8.0 drops ImageMagick and speeds up image-to-PDF conversion up to 7.2× — llama_index · 2026-07-25
- Claude Opus 5 flips between 1/3 and 2/5 before finally settling on an answer — Sauers_ · 2026-07-25
- Claude Opus 5 rates its own moral patienthood at 41% in automated interviews — Sauers_ · 2026-07-25
- DeepSeek deprecates `deepseek-chat` and `deepseek-reasoner` for V4-Flash and V4-Pro — thejasminejade · 2026-07-25