Project APE launches CRED to test whether LLMs can verify research errors
soumitrashukla9 · x · 2026-07-22
Project APE introduces CRED for testing autonomous verification of research errors
The post shares a new paper, Verifying the Verifiers: Towards Autonomous Policy Evaluation, and a new benchmark called CRED.
- The authors ask whether LLMs can automatically verify key errors in policy-evaluation research as generation of plausible papers becomes cheap.
- They evaluate 60+ LLMs on the benchmark.
- The accompanying chart shows a cost-recall frontier across model release dates, comparing systems such as Mistral NeMo 12B, DeepSeek V4 Pro/Flash, Hunyuan 3, Kimi K2.6, GLM 5.2, and Codex 5.6 Luna high.
- The broader goal of Project APE is to build a system that can learn to do policy evaluation autonomously, making verification cheaper and more reliable at scale.
Related event: Project APE Evaluates LLMs as Autonomous Research Error Verifiers(5 posts)→
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11