Anthropic's internal bar: models need 85% to replace researchers; Mythos 5.1 far off

haider1 · x · 2026-09-02

Per Anthropic's internal evaluation info, Mythos 5.1 scored slightly below Opus 5 but above Mythos 5 on the company's internal real-R&D benchmark.

Anthropic believes a model would need 85% to genuinely replace its research staff — and Mythos 5.1 is still nowhere near that bar, a rare quantified internal reference for how far frontier models are from replacing AI researchers.

Original post →

More from Models

Models channel →