PG-LLM benchmark finds Claude Opus 5 tops 217 protein-variant tasks
LeoTZ03 · x · 2026-07-28
PG-LLM benchmark for protein variant effect prediction
PG-LLM introduces a benchmark for testing whether general-purpose LLMs can predict protein variant effects.
- The benchmark spans 217 tasks.
- In the reported ranking of 50 variants, Claude Opus 5 (Max) leads the tested LLMs with Spearman ρ = 0.406.
- It also outperforms 49 of 95 specialized protein predictors in that setting.
The chart compares several models, including GPT-5.x, Gemini, Kimi, and GLM variants, against sequence-only and alignment/structure-based protein methods, with median reference lines shown for the two groups.
Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23