Claim Verification Benchmarks Mostly Test Retrieval, Not Reasoning, Finds 24K-Trace Study
deliprao · x · 2026-10-06
A new arXiv paper by Delip Rao and Chris Callison-Burch analyzes reasoning traces generated with GPT-4o-mini for 24K claim-verification examples across 9 datasets:
- Direct evidence extraction dominates, while multi-sentence synthesis and numerical reasoning are severely under-represented — high benchmark scores mainly reflect retrieval-plus-entailment ability.
- Dataset-level biases are stark: some datasets almost exclusively test lexical matching; others require synthesis in roughly half of cases.
- Using a compact 1B-parameter reasoning verifier, the authors characterize five error types: general-domain verification suffers from lexical-overlap bias, scientific verification from overcautiousness, and mathematical verification from arithmetic failures.
- The paper outlines recommendations for harder evaluation suites that better test real reasoning.
The post also celebrates PhD student Maxine Liu's first UPenn NLP paper and invites feedback at the poster session.
More from Models
- Opus 5.5 is efficient on subscription, not via API — 6.1 remains the workhorse — haider1 · 2026-10-06
- Hiding Y-Axis Labels in Early nanogpt Benchmarks Is "Academic Dishonesty" — PMinervini · 2026-10-06
- LLM MoEs run at ~5% sparsity, cited as counterexample in consciousness complexity debate — JoshPurtell · 2026-10-06
- Report: Zhipu's GLM 5.3 Also Hit a Delayed Release — teortaxesTex · 2026-10-06
- Is There a Market for the 10th-Best Open Model? $5B Capex Question Sparks Debate — ericjang11 · 2026-10-06
- How AA Benchmarks 26 Search API Products Across 13 Providers — ArtificialAnlys · 2026-10-06