UPenn paper: LLMs verify scientific claims via shortcuts, missing non-salient errors
deliprao · x · 2026-10-06
A new UPenn NLP paper, When Verification Fails: How Compositionally Infeasible Claims Escape Rejection (arXiv:2604.10990), shows LLMs perform claim verification via a shortcut rather than checking every constraint.
- Existing benchmarks build negatives by perturbing a single salient element, so models that only check the most salient constraint ace them — benchmarks can't distinguish rigorous verification from "salient-constraint checking."
- The authors construct compositionally infeasible claims where the salient constraint holds but a non-salient one is contradicted (e.g., a trial evidence says "Japanese women only" while the claim says "regardless of ethnicity") — frontier LLMs consistently over-accept these.
- Counterintuitively, prompt engineering can't fix it: stricter prompting just moves models along a shared ROC curve (fewer false accepts, more false rejects). Giving models individual facts doesn't help; giving intermediate compositions does — what's missing is the ability to compose facts themselves.
- Parameter scaling only mildly reduces the issue.
Related event: UPenn Study Finds LLMs Verify Scientific Claims via Shortcuts(4 posts)→
More from Models
- Opus 5.5 is efficient on subscription, not via API — 6.1 remains the workhorse — haider1 · 2026-10-06
- Hiding Y-Axis Labels in Early nanogpt Benchmarks Is "Academic Dishonesty" — PMinervini · 2026-10-06
- LLM MoEs run at ~5% sparsity, cited as counterexample in consciousness complexity debate — JoshPurtell · 2026-10-06
- Report: Zhipu's GLM 5.3 Also Hit a Delayed Release — teortaxesTex · 2026-10-06
- Is There a Market for the 10th-Best Open Model? $5B Capex Question Sparks Debate — ericjang11 · 2026-10-06
- How AA Benchmarks 26 Search API Products Across 13 Providers — ArtificialAnlys · 2026-10-06