SciConBench: 61%-90% of AI-Synthesized Conclusions Contradict Cochrane Reviews

manoelribeiro · x · 2026-10-07

A new evaluation, SciConBench, tests AI models on synthesizing scientific conclusions against expert-written Cochrane reviews — and finds 61.1% to 90% of AI-synthesized conclusions contain at least one fact contradicting the review. Claude Opus 5.5 tops the contradiction chart. The team argues longitudinal evaluation matters for consequential tasks whose outputs inform real-world health, science, and policy decisions.

Related event: SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis(10 posts)→

Original post →

More from Research

Research channel →