ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark
LeoTZ03 · x · 2026-07-28
A new ProteinGym-LLM study benchmarks general-purpose LLMs on a protein-variant ranking task and finds that Claude Opus 5 (Max) leads the tested models on the reported subset.
Key points:
- The benchmark asks a model to rank 50 protein variants from highest to lowest experimental fitness.
- It uses Spearman ρ to compare model ordering against experimental results.
- Across 217 tasks, Claude Opus 5 (Max) reaches ρ = 0.406.
- The authors say it outperforms 49 of 95 specialized protein predictors when ranking 50 variants.
The article also explains the evaluation setup: the model gets only an assay description, wild-type sequence, and shuffled mutants — no fitness labels, examples, alignment, or structure.
Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→
More from Models
- Bug Hunt Bench ranks GPT-6 Astra top as coding models fix real planted bugs, costs spread 200x — PawelHuryn · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23