METR reportedly used Redwood's conceptual reasoning benchmark to eval Mythos 5.1
dfrsrchtwts · x · 2026-09-02
A tweet notes that METR apparently used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1, and that METR published more text on Mythos 5.1 than seen in past model cards. Third-party observation, unconfirmed.
More from Models
- Claude Fable 5.1 cuts cache read costs by 75% — rohanpaul_ai · 2026-09-02
- Anthropic releases Claude Fable 5.1 with better performance and lower cost — dr_cintas · 2026-09-02
- Perplexity adds Claude Fable 5.1, cutting costs by 37% — perplexity_ai · 2026-09-02
- Anthropic Fable 5.1 System Prompt Leaked, Spanning 270k+ Characters — Scobleizer · 2026-09-02
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Multi-agent evals lack model comparisons, need more details — scaling01 · 2026-09-02