SciConBench: 61%-90% of AI-Synthesized Conclusions Contradict Cochrane Reviews
manoelribeiro · x · 2026-10-07
A new evaluation, SciConBench, tests AI models on synthesizing scientific conclusions against expert-written Cochrane reviews — and finds 61.1% to 90% of AI-synthesized conclusions contain at least one fact contradicting the review. Claude Opus 5.5 tops the contradiction chart. The team argues longitudinal evaluation matters for consequential tasks whose outputs inform real-world health, science, and policy decisions.
More from Research
- Scott Alexander's open letter to Steven Pinker: g-factor is real and AI scaling will keep climbing — Astral Codex Ten · 2026-10-07
- Paper shows diffusion transformer tokens encode lots of image info before it's interpretable — kwangmoo_yi · 2026-10-07
- Why synthetic cells die after five generations: they can't recycle their own broken parts — NikoMcCarty · 2026-10-07
- Survey: The Numerical Linear Algebra behind Large Language Models — Abdelkader Baggag · 2026-10-07
- Harvard releases privacy-preserving agent tool for blood biomarker discovery, +4.18 AUC points median — Harvard · 2026-10-07
- Stanford paper: one agent per user underperforms a single shared agent on common resources — rohanpaul_ai · 2026-10-07