61–90% of AI-synthesized medical conclusions contain factual errors, SciConBench finds
manoelribeiro · x · 2026-10-07
A dynamic benchmark, SciConBench, continuously evaluates frontier models on scientific conclusion synthesis using SciConHarness, which blocks the ground-truth Cochrane review and other answer-revealing sources to force genuine synthesis rather than lookup.
- 61.1%–90% of AI-synthesized conclusions contain at least one fact contradicting the expert-written Cochrane review, with Claude Opus 5.5 topping the contradiction chart.
- Synthesis is barely improving: even the best model (GPT-6.1 Sol) reaches only 41.1 factual F1, with most models clustered around 0.30–0.33.
- The framework combines a fixed core set with a fresh monthly rolling panel of newly published reviews to separate genuine gains from leakage, and could extend beyond healthcare.
- The team plans to run the benchmark every two months for 1–2 years and is seeking API access, evaluation credits, or funding to sustain the public leaderboard.
More from Models
- Token-based pricing is strange: unpredictable, decoupled from value, misaligned incentives — amankhan · 2026-10-07
- JevBench splits leaderboard: open-weight models and API providers now ranked separately — airesearch12 · 2026-10-07
- Hot take: active params barely matter for cyber capability evals — RL env coverage is key — teortaxesTex · 2026-10-07
- Decagon launches Voice 3 with Chord voice model and duplex architecture for customer agents — Scobleizer · 2026-10-07
- JEV-9B, a Qwen3.5-based calibrated decision model, trends on Hugging Face — autotrust · 2026-10-07
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07