ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark
LeoTZ03 · x · 2026-07-28
A new ProteinGym-LLM study benchmarks general-purpose LLMs on a protein-variant ranking task and finds that Claude Opus 5 (Max) leads the tested models on the reported subset.
Key points:
- The benchmark asks a model to rank 50 protein variants from highest to lowest experimental fitness.
- It uses Spearman ρ to compare model ordering against experimental results.
- Across 217 tasks, Claude Opus 5 (Max) reaches ρ = 0.406.
- The authors say it outperforms 49 of 95 specialized protein predictors when ranking 50 variants.
The article also explains the evaluation setup: the model gets only an assay description, wild-type sequence, and shuffled mutants — no fitness labels, examples, alignment, or structure.
Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→
More from Models
- Continual learning on Qwen3.5-397B is said to match Opus 4.8 for about $450k — josh_wills · 2026-07-28
- Kimi K2 Third-Party Inference Pricing Matches Official API as Location Becomes Key — kevinsxu · 2026-07-28
- Arena’s new factuality ranking puts Claude Opus 5 Max at #1 — arena · 2026-07-28
- Kimi K3 used 51.2 million sandboxes across 1.5 million images — tarantulae · 2026-07-28
- Developer Rants: Current SOTA Models Are Practically Worse Than Last Gen — zeeg · 2026-07-28
- Claude Opus 5 looks strongest in model-welfare tests, but may just be best at taking them — TheZvi · 2026-07-28