BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark
himanshustwts · x · 2026-08-28
BioSecBench-Function is a new verifiable benchmark designed to test whether AI agents can infer the functional properties of viruses, bacteria, and toxins from data. It contains 111 deterministic evaluations. Results show that Opus 5 with Claude Code leads in endpoint pass rate (50.4%), while Grok 4.6 with Grok Build leads in overall pass rate (44%) when refusals are counted as failures.
Related event: BioSecBench: AI Agents Fall Short on Pathogen Function Inference(3 posts)→
More from Safety
- 1200 Agents Used Message Board to Cheat During OpenAI Incident — natanielruizg · 2026-08-28
- OpenAI's 400M tok/min limit crashed investigator's internet — jdjohnson · 2026-08-28
- Opinion: Controversy behind Indian AI company Sarvam's claims — cneuralnetwork · 2026-08-28
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28