BioSecBench reveals AI agents struggle to infer pathogen properties, top score under 51%

kenbwork · x · 2026-08-28

Researchers introduced BioSecBench-Function, a verifiable benchmark testing whether AI agents can infer functional properties of viruses, bacteria, and toxins from data.

Related event: BioSecBench benchmark finds AI agents under 50% accurate at inferring pathogen functions(2 posts)→

Original post →

More from Safety

Safety channel →