Paper proposes a CRED taxonomy and benchmark to measure research-error detectors

soumitrashukla9 · x · 2026-07-22

The paper proposes three ingredients for measuring verifier quality in AI research: a versioned CRED taxonomy for research errors, a benchmark with deterministic scoring to quantify detection of real errors, and a longitudinal evaluation of verifiers, including both models and harnesses. The image adds concrete error classes such as prose-table inconsistency, paper-code inconsistency, execution/reproducibility failures, and data-integrity issues.

Related event: Project APE Introduces CRED Benchmark for LLM Research Error Verification(5 posts)→

Original post →

More from Research

Research channel →