Audit of 11 agentic AI systems finds they still fail at scientific evidence synthesis
manoelribeiro · x · 2026-07-29
A new audit finds that agentic AI systems struggle to synthesize scientific evidence reliably.
- The authors evaluated 11 AI and agentic systems, including proprietary research agents and clinical tools like OpenEvidence.
- They used questions derived from the Cochrane Database of Systematic Reviews.
- Systems were prompted to generate conclusions from retrieved evidence and then compared with expert-written answers.
- The main failure modes were unsupported claims, contradictory conclusions, and missed key facts.
- The paper argues these systems can still propagate misleading medical or scientific conclusions, so they need much stricter auditing before high-stakes use.
The takeaway is not that search-based AI is useless, but that current agentic setups are not yet trustworthy as evidence synthesizers.
More from Research
- Anthropic’s Claude Mythos AI reportedly finds new attacks on cryptographic algorithms — luisdans · 2026-07-29
- OKLS brings KL-optimal Shampoo to language model training with 1.45× parameter efficiency — aryaman2020 · 2026-07-29
- Blog Post Explains Why Tokenizer-Free Language Modeling Isn't Actually Tokenizer-Free — cjmaddison · 2026-07-29
- Fish Audio open-sources S2, a controllable TTS model trained on 10M hours of audio — alex_verem · 2026-07-29
- NVIDIA shares a guide to training open models with RL on Prime Intellect — NVIDIAAI · 2026-07-29
- AIR-seq tracks RNA synthesis and decay using standard RNA-seq libraries — lpachter · 2026-07-29