Claude Opus 5 tops FrontierCode at 53.4% and 63.6% with medium test-time compute

andrew_n_carr · x · 2026-07-25

A FrontierCode chart shows Claude Opus 5 scaling strongly with more test-time compute on both the main and extended sets.

The post uses the benchmark to argue that Opus 5 is a very strong coding model, especially when given the right reasoning budget.

Related event: Anthropic Releases Claude Opus 5: Comprehensive Performance Leap at Half the Competitor's Price(67 posts)→

Original post →

More from Models

Models channel →