PG-LLM benchmark finds Claude Opus 5 tops 217 protein-variant tasks
LeoTZ03 · x · 2026-07-28
PG-LLM benchmark for protein variant effect prediction
PG-LLM introduces a benchmark for testing whether general-purpose LLMs can predict protein variant effects.
- The benchmark spans 217 tasks.
- In the reported ranking of 50 variants, Claude Opus 5 (Max) leads the tested LLMs with Spearman ρ = 0.406.
- It also outperforms 49 of 95 specialized protein predictors in that setting.
The chart compares several models, including GPT-5.x, Gemini, Kimi, and GLM variants, against sequence-only and alignment/structure-based protein methods, with median reference lines shown for the two groups.
Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→
More from Models
- ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark — LeoTZ03 · 2026-07-28
- Kimi K2 Third-Party Inference Pricing Matches Official API as Location Becomes Key — kevinsxu · 2026-07-28
- Arena’s new factuality ranking puts Claude Opus 5 Max at #1 — arena · 2026-07-28
- Kimi K3 used 51.2 million sandboxes across 1.5 million images — tarantulae · 2026-07-28
- Developer Rants: Current SOTA Models Are Practically Worse Than Last Gen — zeeg · 2026-07-28
- A parody chart claims GPT wins because its version number is larger — Tystros · 2026-07-28