Claude Opus 5 tops two new biology benchmarks at 72.5% and 60.6%
anshulkundaje · x · 2026-07-25
- Anthropic’s Claude Opus 5 is being pitched as a model that comes close to the frontier intelligence of Fable 5 at half the price.
- The attached system-card excerpt highlights two new biology benchmarks built by LatchBio:
- SpatialBench Verified: 115 externally validated real-world spatial transcriptomics problems.
- SingleCellBench Verified: 195 problems covering common single-cell RNA-seq workflows such as cell labeling, differential expression, and batch correction.
- On SpatialBench Verified, Claude Opus 5 scores 72.5%, ahead of Claude Mythos 5 (69.2%), Claude Sonnet 5 (67.8%), and Claude Opus 4.8 (66.6%).
- On SingleCellBench, Opus 5 leads again at 60.6%, versus Mythos 5 (59.3%), Opus 4.8 (58.2%), and Sonnet 5 (56.2%).
- The surrounding commentary says frontier biology benchmarking now requires close scientist-engineer collaboration and no longer has a settled playbook.
Related event: Anthropic Launches Claude Opus 5 with SOTA Coding Performance(80 posts)→
More from Models
- Anthropic’s latest chart crime looks like over-trusting Claude, not deliberate hype — herbiebradley · 2026-07-25
- Claude Opus 5 fixes 11 of 45 hidden bugs, versus 2 for Opus 4.8 — PawelHuryn · 2026-07-25
- Miles Brundage says Anthropic’s chart issue is really about over-trusting Claude — Miles_Brundage · 2026-07-25
- A benchmark chart becomes an AI meme after viewers spot the messy numbers — Miles_Brundage · 2026-07-25
- FrontierCode 1.1 shows Opus 5 can score lower under stricter reasoning settings — andrew_n_carr · 2026-07-25
- Bug Hunt Bench: GPT-5.6 Sol fixes 22 bugs, Opus 5 12, on a 45-bug repo — PawelHuryn · 2026-07-25