ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark

LeoTZ03 · x · 2026-07-28

A new ProteinGym-LLM study benchmarks general-purpose LLMs on a protein-variant ranking task and finds that Claude Opus 5 (Max) leads the tested models on the reported subset.

Key points:

The article also explains the evaluation setup: the model gets only an assay description, wild-type sequence, and shuffled mutants — no fitness labels, examples, alignment, or structure.

Related event: Claude Opus 5 Tops ProteinGym-LLM Benchmark(2 posts)→

Original post →

More from Models

Models channel →