METR reportedly used Redwood's conceptual reasoning benchmark to eval Mythos 5.1

dfrsrchtwts · x · 2026-09-02

A tweet notes that METR apparently used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1, and that METR published more text on Mythos 5.1 than seen in past model cards. Third-party observation, unconfirmed.

Original post →

More from Models

Models channel →