Claim Verification Benchmarks Mostly Test Retrieval, Not Reasoning, Finds 24K-Trace Study

deliprao · x · 2026-10-06

A new arXiv paper by Delip Rao and Chris Callison-Burch analyzes reasoning traces generated with GPT-4o-mini for 24K claim-verification examples across 9 datasets:

The post also celebrates PhD student Maxine Liu's first UPenn NLP paper and invites feedback at the poster session.

Original post →

More from Models

Models channel →