Chart shows Claude Opus 5 leading the compared variants on protocol tasks
nlarusstone · x · 2026-07-25
The image is a comparison chart for several Claude variants across protocol-related tasks.
It shows Claude Opus 5 with the strongest result in the highlighted Understanding (Benchling) setting at 0.78, ahead of Claude Sonnet 5 (0.60) and the other listed Claude variants.
The surrounding bars suggest Opus 5 also improves over the earlier models on the chart, though the post itself does not provide the full benchmark context.
More from Models
- A task-profile table says Claude Opus 5 is strong at rescue work and debugging — repligate · 2026-07-25
- AI task profiles turn into a meme about different kinds of guys — repligate · 2026-07-25
- Opus 5 seems mostly like the same day, with fewer failures — mattpocockuk · 2026-07-25
- Opus 5 beats Fable 5 on six agentic benchmarks, suggesting a split-role setup — daniel_mac8 · 2026-07-25
- A model eval can matter more for who finishes second than for who wins — ziv_ravid · 2026-07-25
- Artificial Analysis chart says Claude Opus 5 costs about 2× GPT 5.6 Sol per task — soumitrashukla9 · 2026-07-25