Benchmarking LLMs on false closure failures
Plastic-Cell-4497 · reddit · 2026-08-31
The author built a benchmark called CFC targeting a specific LLM failure mode: False Closure, where models reach definite conclusions unjustified by evidence. This includes silently turning UNRESOLVED to TRUE/FALSE, resolving conflicts without rules, and relying on stale evidence.
The test suite includes 100 cases run across four model tracks (1,200 runs). Strong models like Claude and Gemini passed 98.7% semantically, but failed on different cases. The GPT track scored perfectly but was marked as non-independent. The focus is not on the score gap but on the specific reasoning errors that survive high pass rates, inviting feedback from the evaluation community.
More from Research
- Chollet responds to ARC-AGI eval dispute: don't claim untested scores — fchollet · 2026-08-31
- Research复盘:Linear attention found ineffective in specific setup — jm_alexia · 2026-08-31
- Google releases GlucoFM, a lightweight foundation model for improved metabolic predictions — thione · 2026-08-31
- LAION releases 10M-hour open video dataset LAION-BVD for multimodal pre-training — thione · 2026-08-31
- Figure launches Index dataset, claiming world's largest and most diverse robot training data — thione · 2026-08-31
- NYU Paper: Scalable, Personalized Oral Assessments Using Voice AI for Under $1 per Exam — ipeirotis · 2026-08-31