Project APE launches CRED to test whether LLMs can verify research errors
soumitrashukla9 · x · 2026-07-22
Project APE introduces CRED for testing autonomous verification of research errors
The post shares a new paper, Verifying the Verifiers: Towards Autonomous Policy Evaluation, and a new benchmark called CRED.
- The authors ask whether LLMs can automatically verify key errors in policy-evaluation research as generation of plausible papers becomes cheap.
- They evaluate 60+ LLMs on the benchmark.
- The accompanying chart shows a cost-recall frontier across model release dates, comparing systems such as Mistral NeMo 12B, DeepSeek V4 Pro/Flash, Hunyuan 3, Kimi K2.6, GLM 5.2, and Codex 5.6 Luna high.
- The broader goal of Project APE is to build a system that can learn to do policy evaluation autonomously, making verification cheaper and more reliable at scale.
Related event: Project APE Introduces CRED to Test LLM Error Verification(2 posts)→
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22