Redwood and Anthropic release the Conceptual Reasoning Index (CRI)
RyanGreenblatt · x · 2026-08-13
Redwood Research, in collaboration with Anthropic, has released the Conceptual Reasoning Index (CRI) to address the challenge of measuring and automating AI safety research. While AI excels in domains with cheap and reliable feedback like math and coding, evaluating complex safety evidence—such as whether an experiment predicts the behavior of much smarter future models—relies heavily on sound argumentation. Fully manually verified by researchers, the CRI dataset is designed to evaluate models' reasoning abilities on hard-to-adjudicate safety questions, aiming to steer away from the risks of unintelligibly automating AI R&D.
Related event: Redwood and Anthropic Introduce Conceptual Reasoning Index(2 posts)→
More from Safety
- Surge in Cyberattacks May Slow Enterprise Momentum and Reinforce Single-Founder Model — natesiggard · 2026-08-13
- Report: AI Data Centers Consume 10x Indirect Water, Competing with Western Farmers — AndyMasley · 2026-08-13
- Twitch Opts Streamers In by Default to Train Amazon's AI Models — 404 Media · 2026-08-13
- Hinton, Fei-Fei Li and Andrew Ng Debate AI Regulation and Open Source — TechCrunch AI · 2026-08-13
- Far-Right Imageboard Users Building 'Fully Autonomous AI Doxxing' Tools — joeddav · 2026-08-13
- Implementing AI Text Watermarks via Constrained Sampling for EU AI Act Compliance — Xianbao_QIAN · 2026-08-13