Claude Opus 5 looks especially strong on visually grounded benchmarks

xeophon · x · 2026-07-25

A reply post cites Claude Opus 5’s benchmark chart and summarizes the result as especially strong on tasks with a visual component.

The referenced chart shows Opus 5 ahead on several benchmarks such as Frontier-Bench terminal coding, GDPVal-AA v2, ARC-AGI-3, BrowseComp, OSWorld 2.0, and BioMysteryBench. The commenter’s angle is that the model looks particularly strong wherever visuals are part of the task.

Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(32 posts)→

Original post →

More from Models

Models channel →