Opus 5 tops nine biology benchmarks, but still trails in some analysis tasks
kenbwork · x · 2026-07-25
A benchmark run on Opus 5 across nine agentic biology tasks says the model is the best evaluated so far on several areas, including variant discovery, preclinical pharmacology, genomic surveillance, and long-horizon single-cell analysis.
Key results
- Variant discovery and interpretation: +7.6 pp over Sol 5.6
- Small-molecule preclinical pharmacology: +7.0 pp over Sol
- Genomic surveillance: +7.3 pp over Sol
- Long-horizon single-cell analysis: +3.2 pp over Sol
The analysis says Opus 5 beats every previous Anthropic model on the suite, taking 6 of 7 benchmarks from Opus 4.8, often by 8–16 points. But it is not uniformly better:
- On short-horizon single-cell analysis, it trails Sol 5.6 by 2 points (60.1% vs 62.1%).
- On long-horizon spatial biology, it more than doubles Opus 4.8 (25.0% vs 11.1%) but still loses to Sol (38.9%).
- It also regresses on epigenomics analysis, where Opus 4.8 is 4.2 points higher.
The post is a continuation of a larger article on how good Opus 5 is at biology, with the benchmark dashboard now live on benchmarks.bio.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11