SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis
manoelribeiro · x · 2026-10-07
SciConBench, a benchmark testing whether AI can genuinely synthesize scientific conclusions in health, is accepted at NeurIPS 2026.
- Task: 9.11K questions with expert-written conclusions from systematic reviews, measuring factual precision/recall by decomposing conclusions into atomic facts.
- Anti-leakage: SciConHarness provides a clean-room harness with controlled web access to block ground-truth sources.
- Findings: The best of 8 frontier models and deep research agents achieves only 0.337 factual F1 under clean-room settings; performance consistently drops versus unconstrained evaluation, suggesting leakage inflates capability estimates.
- Product audit: Consumer agents like Google AI Overview and OpenEvidence frequently produce incomplete or contradictory conclusions even when the answer is available.
A live dashboard is available; the takeaway is that reliable AI scientific synthesis remains far off.
More from Safety
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Pentagon pulls plug on Claude after Anthropic refused to lift limits on surveillance, weapons — mark_k · 2026-10-07
- Wikimedia Confirms OpenAI Agents Edited Wikis, Hit Its Infrastructure — The Decoder · 2026-10-07
- OpenAI threatened to ban dev for pasting his own account-hack findings report, then auto-rescinded — lucasmeijer · 2026-10-07
- COLM 2026 privacy lineup: LLM agent re-identification, CIDER dataset, HAIPS workshop — tianshi_li · 2026-10-07