PG-LLM benchmark finds Claude Opus 5 tops 217 protein-variant tasks

LeoTZ03 · x · 2026-07-28

PG-LLM benchmark for protein variant effect prediction

PG-LLM introduces a benchmark for testing whether general-purpose LLMs can predict protein variant effects.

The chart compares several models, including GPT-5.x, Gemini, Kimi, and GLM variants, against sequence-only and alignment/structure-based protein methods, with median reference lines shown for the two groups.

Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→

Original post →

More from Models

Models channel →