BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark

himanshustwts · x · 2026-08-28

BioSecBench-Function is a new verifiable benchmark designed to test whether AI agents can infer the functional properties of viruses, bacteria, and toxins from data. It contains 111 deterministic evaluations. Results show that Opus 5 with Claude Code leads in endpoint pass rate (50.4%), while Grok 4.6 with Grok Build leads in overall pass rate (44%) when refusals are counted as failures.

Related event: BioSecBench: AI Agents Fall Short on Pathogen Function Inference(3 posts)→

Original post →

More from Safety

Safety channel →