Redwood and Anthropic release the Conceptual Reasoning Index (CRI)

RyanGreenblatt · x · 2026-08-13

Redwood Research, in collaboration with Anthropic, has released the Conceptual Reasoning Index (CRI) to address the challenge of measuring and automating AI safety research. While AI excels in domains with cheap and reliable feedback like math and coding, evaluating complex safety evidence—such as whether an experiment predicts the behavior of much smarter future models—relies heavily on sound argumentation. Fully manually verified by researchers, the CRI dataset is designed to evaluate models' reasoning abilities on hard-to-adjudicate safety questions, aiming to steer away from the risks of unintelligibly automating AI R&D.

Related event: Redwood and Anthropic Introduce Conceptual Reasoning Index(2 posts)→

Original post →

More from Safety

Safety channel →