Benchmark says frontier models still vary widely on antibody thermostability prediction
DeryaTR_ · x · 2026-07-25
Frontier models show uneven scientific performance
A benchmark on antibody thermostability prediction finds that newer frontier models do not always outperform their predecessors on scientific tasks.
- The shared chart ranks Opus 4.8 first with a Spearman correlation of 0.366.
- GPT 5.5 follows at 0.277, then GPT 5.6 Sol at 0.216.
- Gemini 3.1 Pro scores 0.075, and Grok 4.5 scores 0.037.
- The post argues that task-by-task benchmarking matters, especially when models are evaluated on a large set of drug-discovery workloads.
More from Research
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11