Opus 5 ARC-AGI-3 达到 30.2%,多项代理基准领先
socoolandawesome · reddit · 2026-07-25
Opus 5 多项基准领先,ARC-AGI-3 达到 30.2%
配图展示了 Opus 5 在多项基准上的成绩,包括:
- Agentic Terminal Coding:43.3%
- Knowledge Work:1861
- ARC-AGI-3:30.2%
- Agentic Search:90.8%
- Humanity’s Last Exam:无工具 56.3%,有工具 64.7%
- OSWorld 2.0:70.6%
- DeepSWE v1.1:68.8%
- FrontierCode v1.1 Main:53.4%
图表将它与 Fable 5、Opus 4.8 和 GPT-5.6 Sol 对比,整体结论是:Opus 5 在代理式编码、搜索和新问题解决上表现更强。
所属事件:Anthropic 发布 Claude Opus 5(35 条相关)→
「模型」频道最新
- Anthropic 推出 Claude Opus 5,定价仍是 $5/$25 — ctjlewis · 2026-07-25
- Claude Opus 5 在提示注入鲁棒性榜上接近满分 — EricBuess · 2026-07-25
- 用户调侃 Claude Opus 5 像 GPT-5,Anthropic 则强调 SOTA 成绩 — ChengleiSi · 2026-07-25
- Claude Opus 5 编程好用,但把 Compound Engineering 搞坏了 — danshipper · 2026-07-25
- Claude Opus 5 被描述为单任务更便宜、表现却很强 — EricBuess · 2026-07-25
- Anthropic 发布 Claude Opus 5,主打生物和化学能力 — gabepgomes · 2026-07-25