Anthropic Internal Eval: Mythos 5.1 Scores Lower Than Opus 5
scaling01 · x · 2026-09-02
Anthropic's internal CoBench evaluation shows that the new Mythos 5.1 model scores lower than Opus 5.
More from Models
- Anthropic Fable 5.1 System Prompt Leaked, Spanning 270k+ Characters — Scobleizer · 2026-09-02
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Multi-agent evals lack model comparisons, need more details — scaling01 · 2026-09-02
- Fable 5.1 quietly watermarks all generated text — adityavg13 · 2026-09-02
- Fable 5.1 Beats GPT-5.6 on Benchmark at Lower Cost — haider1 · 2026-09-02
- Claude Fable 5.1 adds heavy instructions,疑似过度对齐 — teodorio · 2026-09-02