Opus 5 launches with strong scores on agentic coding, search, and computer use
TheInfiniteUniverse_ · reddit · 2026-07-25
A post announces that Opus 5 has just dropped, with a benchmark table showing its performance across agentic coding, search, reasoning, computer use, business workflows, legal, health, and biology. The image suggests strong results in agentic terminal coding, knowledge work, agentic search, computer use, and several domain-specific benchmarks, while comparing Opus 5 against Fable 5, Opus 4.8, and GPT-5.6 Sol.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11