Opus 5 tops nine biology benchmarks, but still trails in some analysis tasks

kenbwork · x · 2026-07-25

A benchmark run on Opus 5 across nine agentic biology tasks says the model is the best evaluated so far on several areas, including variant discovery, preclinical pharmacology, genomic surveillance, and long-horizon single-cell analysis.

Key results

The analysis says Opus 5 beats every previous Anthropic model on the suite, taking 6 of 7 benchmarks from Opus 4.8, often by 8–16 points. But it is not uniformly better:

The post is a continuation of a larger article on how good Opus 5 is at biology, with the benchmark dashboard now live on benchmarks.bio.

Original post →

More from Models

Models channel →