BioSecBench: AI Agents Fall Short on Pathogen Function Inference
Researchers introduced BioSecBench-Function, a benchmark of 111 deterministic evaluations testing whether AI agents can infer functional properties of pathogens; results show overall accuracy below 50%, with Opus and Grok ranking top two.
2026-08-28 ~ 2026-08-28 · 3 related posts
- BioSecBench reveals AI agents struggle to infer pathogen properties, top score under 51% — kenbwork · 2026-08-28
- AI Agents to Bridge Pathogen Characterization Gap: New Benchmarks Reveal Current Limits — kenbwork · 2026-08-28
- BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark — himanshustwts · 2026-08-28