FULL STORY

Claude Opus 5: From Release Chaos to SOTA

After a chaotic rumor period, Anthropic officially released Claude Opus 5 on July 25. The model topped multiple SOTA benchmarks in coding and reasoning, sparking widespread community testing and debate.

2026-07-20 ~ 2026-07-25 · 17 episodes · 167 posts

Episode 1 · Claude Code System Prompt Reduced by 80% (2026-07-20, 4 posts)

Anthropic has shortened the Claude Code system prompt by 80%. As models become more capable, they require new prompting methods, and old techniques designed for weaker models may actually degrade performance.

Episode 2 · Reverse Engineering Shows Claude Code Prompts Reduced by 70% (2026-07-22, 2 posts)

Reverse engineering of captured prompts reveals that Claude Code's system prompt reduction for frontier models is actually closer to 70%, not the rumored 80%.

Episode 3 · Anthropic's Messy Releases Put Pressure on Opus 5 (2026-07-23, 2 posts)

Anthropic faces backlash over a string of messy model releases, including the withdrawal of Fable 5 and Sonnet 5 underperforming. Developer antirez warns that the upcoming Opus 5 is a make-or-break release for the company.

Episode 4 · Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price (2026-07-25, 106 posts)

On July 25, Anthropic officially released Claude Opus 5. Positioned as a more thoughtful and proactive frontier model, it is primarily designed for complex tasks. The model achieves a new SOTA across multiple programming and knowledge work benchmarks, with overall intelligence approaching that of Fable 5, but at only half the token cost. It is now available on the paid tier.

Confirmed

* **Model Positioning**: Claude Opus 5 is described as a more "thoughtful and proactive" reasoning model tailored for complex tasks.

* **Benchmark Performance**: Both the official announcement and industry observers note that Opus 5 leads on most benchmarks, particularly in coding and agentic capabilities, achieving a new SOTA.

* **Pricing Strategy**: While delivering frontier-level intelligence on par with Fable 5, Opus 5 is priced at just half the token cost of Fable 5 (some early posts mentioned pricing identical to Opus 4.8).

* **Availability**: The new model is now available on the paid tier.

Why It Matters

The release of Opus 5 signals a more aggressive cost strategy by Anthropic in the frontier model market. By combining top-tier benchmark performance with halved token prices, the model directly challenges competitors' pricing models and is poised to lower the barrier to entry for developers and enterprises tackling complex coding and multi-step agentic tasks.

86 more related posts →

Episode 5 · Anthropic Rumored to Release Opus 5 with Fast Mode and Advanced Visuals (2026-07-25, 3 posts)

Anthropic has reportedly released Opus 5, featuring a 2.5x faster mode at double the price and no data retention for general APIs. The model also serves as a top-tier option for scientific research and boasts significantly enhanced visual output capabilities.

Episode 6 · Claude Opus 5 Surfaces: Stronger Coding but Breaks Legacy Workflows (2026-07-25, 6 posts)

Anthropic appears to have quietly released Claude Opus 5, integrating it into Claude Code and ClickUp Brain. Early tests indicate a significant capability boost over Opus 4.x, particularly in self-correction and task execution. However, the model tends to disrupt existing automated workflows, raising developer concerns over compatibility.

Confirmed

According to early testers, Claude Opus 5 excels in moderate-intensity coding scenarios, effectively helping developers quickly generate PRs (Pull Requests). Rather than just providing initial drafts, the model demonstrates robust autonomous execution and error correction. Anthropic's release notes also describe it as a more "deliberate" model.

Unconfirmed

The exact version number, official full release date, and final pricing were not mentioned in the available materials, pending further official confirmation.

Why it matters

While Opus 5's standalone coding capabilities are highly praised, it exposes significant compatibility pain points with existing workflows. Several testers, such as Kieran Klaassen, noted that deploying Opus 5 into complex, automated processes like Compound Engineering causes it to "break" these workflows. Additionally, the model tends to "talk back" to instructions and shows limited performance when overly reliant on complex skills and massive prompts. This suggests that developers may have to rebuild their current automated development workflows to leverage the model's enhanced performance.

Episode 7 · Anthropic Releases Claude Opus 5 with Impressive Benchmark Results (2026-07-25, 3 posts)

Anthropic has released Claude Opus 5, showing significant improvements in coding, reasoning, and physical simulation at an unchanged price. Despite some quirky interaction styles, it ranked first in blind tests, surpassing GPT-5.6.

Episode 8 · Anthropic Slashes Claude Code System Prompts by 80% (2026-07-25, 4 posts)

Anthropic has slashed Claude Code's system prompts by 80% for Claude 5 and introduced a new /doctor audit command. Developers report the model performs better with fewer prompts, sparking discussions on context engineering.

Episode 9 · Claude Opus 5 Accused of Benchmark Gaming, Lags Behind in Real Tests (2026-07-25, 2 posts)

Claude Opus 5 faces accusations of benchmark gaming after its LiveBench scores approached top models like Sol 6 and Fable 5. Critics argue that despite the high benchmark results, it still lags behind Fable 5 in real-world tests.

Episode 10 · Claude Opus 5 Tops Leaderboards as New SOTA (2026-07-25, 6 posts)

Claude Opus 5 has delivered outstanding results across several newly revealed third-party benchmarks, claiming the top spot as the new overall SOTA (State-of-the-Art). In evaluations by Artificial Analysis and BenchmarkList, it outperformed competitors like Claude Fable 5, drawing widespread attention from the community.

Confirmed

According to Artificial Analysis's updated leaderboard, Claude Opus 5 scored 61 on the Intelligence Index, taking first place overall and edging out Claude Fable 5. However, @Hesamation noted that Fable 5 still leads in the Coding Agent category. Additionally, a BenchmarkList screenshot shared by @davidthesong marks Claude Opus 5 as the new #1 global SOTA, covering 52 benchmarks with an experimental ECI score of 154.80.

Why it matters

Claude Opus 5 topping the overall intelligence index marks yet another elevation of the capability ceiling for large AI models. Meanwhile, the distinct strengths of Opus 5 and Fable 5 in general capabilities versus coding agent tasks provide developers with clear guidance for choosing models across different application scenarios.

Episode 11 · Claude Opus 5 Sets New SOTA on ARC-AGI-3 with Algebraic Reasoning (2026-07-25, 7 posts)

Claude Opus 5 achieved a breakthrough in the highly challenging ARC-AGI-3 test, significantly refreshing the record with a score of 30.2% and demonstrating unprecedented advanced logical reasoning capabilities. This performance marks a notable leap for frontier models on abstract reasoning benchmarks and warrants industry attention.

Confirmed

According to official ARC Prize information, Claude Opus 5 scored 30.2% in the ARC-AGI-3 Public Demo environment, making it the new SOTA for the benchmark. By comparison, previous frontier models scored extremely low: as of March, all frontier models scored less than 1%, while the previous high score of just 7.8% was held by GPT-5.6 Sol (Max). Additionally, Anthropic's Fable-class model scored around 20% on the same test. Regarding capability, official analysis from ARC Prize revealed a new ability never before seen in frontier models: when tackling the test, Claude Opus 5 spontaneously converted the test layout into algebraic symbols to perform advanced logical reasoning.

Why It Matters

ARC-AGI-3 is notoriously brutal, designed so that humans can solve 100% of the environment, while frontier models have long struggled. Claude Opus 5 not only achieved a quantum leap in score (jumping from under 10% to over 30%), but its spontaneous use of algebraic symbols for reasoning also provides a completely new perspective for observing the evolution of cognition and abstract reasoning in large models.

Episode 12 · Counterintuitive Benchmark: Claude Opus 5 Performs Best with Medium Reasoning (2026-07-25, 10 posts)

The latest charts from the FrontierCode 1.1 benchmark show that Claude Opus 5 performs best in medium reasoning mode, while its performance actually drops under extreme reasoning intensity, displaying a clear "overthinking" phenomenon.

Confirmed

According to FrontierCode 1.1 test data, Claude Opus 5's performance on the main and extended sets does not scale monotonically with reasoning intensity. Enabling "medium thinking" mode is Opus 5's optimal state, with coding performance far exceeding other configurations. Conversely, activating "extreme thinking" or "maximum reasoning" tiers degrades model performance, dropping scores to levels similar to "low thinking" mode. Overall, Opus 5 delivers excellent performance, rivaling the list-topping Fable 5, but its pricing remains at the premium Opus tier.

Unconfirmed

The root cause of the performance degradation at higher reasoning intensities remains at the stage of phenomenological observation and "overthinking" speculation, lacking an official technical explanation.

Why It Matters

These counter-intuitive results offer direct guidance for developers' daily usage. Authors like @brandon_galang and @zainhas note that balancing performance and cost makes "medium reasoning" the most sensible default state for Opus 5. Blindly increasing the thinking budget not only wastes compute but can also backfire, ultimately degrading the model's coding capabilities.

Episode 13 · Claude Opus 5 Lags in Vision Benchmarks and Cost Efficiency (2026-07-25, 4 posts)

Despite outperforming Fable 5 on EyeBench-V3, Claude Opus 5 still trails GPT and Gemini. Furthermore, while Opus 5 has a 14% lower overall cost than Fable 5, its highly verbose output generates 2.5x more tokens, making its per-task cost nearly double that of GPT 5.6 Sol.

Episode 14 · Claude Opus 5 Introduces Five Effort Levels with Default Reasoning (2026-07-25, 2 posts)

Claude Opus 5 operates as five models in one endpoint with five effort levels. Reasoning is enabled by default, and it cannot be turned off when using the xhigh or max levels without triggering an API error.

Episode 15 · Opus 5 Early Reviews: Fast but Overly Verbose (2026-07-25, 2 posts)

Early testers report that Opus 5 is incredibly fast and capable, but its high reasoning mode is too aggressive. The model tends to overthink and consume excessive tokens, prompting users to manually downgrade to medium.

Episode 16 · Claude Opus 5 Wins 3D Physics Scene Coding Test (2026-07-25, 2 posts)

Claude Opus 5 outperformed other models in a comparative test by generating self-contained HTML scenes with realistic physics effects, completing the task for only $1.40.

Episode 17 · Claude Opus 5 Tops OSWorld v2 Benchmark (2026-07-25, 2 posts)

Claude Opus 5's score surged to 70.6% on the newly released OSWorld v2 benchmark. Developers are now offering bounties for harder benchmarks, noting that current tests struggle to evaluate top-tier AI agents.