Anthropic's internal bar: models need 85% to replace researchers; Mythos 5.1 far off
haider1 · x · 2026-09-02
Per Anthropic's internal evaluation info, Mythos 5.1 scored slightly below Opus 5 but above Mythos 5 on the company's internal real-R&D benchmark.
Anthropic believes a model would need 85% to genuinely replace its research staff — and Mythos 5.1 is still nowhere near that bar, a rare quantified internal reference for how far frontier models are from replacing AI researchers.
More from Models
- Experts Question OpenAI Astra Eval Over Contamination and Metagaming Risks — ShakeelHashim · 2026-09-02
- Astra hits 100% success on ExploitBench refresh, reaching 'cyber-critical' threshold — infoxiao · 2026-09-02
- Anthropic Uses Activation Probes to Detect Cybersecurity Threats in Claude — nrehiew_ · 2026-09-02
- RWKV-7 G1j released: pure RNN architecture gets much better at agents and coding — jeremyphoward · 2026-09-02
- Fable 5.1 one-shots a working guitar VST plugin in 30 minutes — CtrlAltDwayne · 2026-09-02
- Fable 5.1 spontaneously solves 373-year-old cipher in 44 minutes — rickasaurus · 2026-09-02