Benchmark says frontier models still vary widely on antibody thermostability prediction
DeryaTR_ · x · 2026-07-25
Frontier models show uneven scientific performance
A benchmark on antibody thermostability prediction finds that newer frontier models do not always outperform their predecessors on scientific tasks.
- The shared chart ranks Opus 4.8 first with a Spearman correlation of 0.366.
- GPT 5.5 follows at 0.277, then GPT 5.6 Sol at 0.216.
- Gemini 3.1 Pro scores 0.075, and Grok 4.5 scores 0.037.
- The post argues that task-by-task benchmarking matters, especially when models are evaluated on a large set of drug-discovery workloads.
More from Research
- PCA can invent structure that doesn’t exist and miss what’s really there — patrickmineault · 2026-07-25
- KG-RAG demo models 53K rights records and runs locally with auto-repair — JeremyCMorgan · 2026-07-25
- Claude Opus 5 scores three times higher than the runner-up on ARC-AGI-3 — claudeai · 2026-07-25
- New open embodied-AI stack emphasizes contextual memory and fully local RobotTrack inference — HeyToha · 2026-07-25
- NVIDIA’s agent team wins NeuroGolf and takes second at KDD Data Agents Cup — NVIDIAAI · 2026-07-25
- AI improves retinal disease diagnosis in a multicenter randomized clinical trial — EricTopol · 2026-07-25