61–90% of AI-synthesized medical conclusions contain factual errors, SciConBench finds

manoelribeiro · x · 2026-10-07

A dynamic benchmark, SciConBench, continuously evaluates frontier models on scientific conclusion synthesis using SciConHarness, which blocks the ground-truth Cochrane review and other answer-revealing sources to force genuine synthesis rather than lookup.

Related event: SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis(10 posts)→

Original post →

More from Models

Models channel →