Claude Opus 5 在代码智能体基准拿到 67 分
Hesamation · x · 2026-07-25
这条转发配图是一张 Artificial Analysis Coding Agent 排行图,比较了多个模型在 DeepSWE、Terminal-Bench v2、SWE-Atlas-QnA 上的综合 pass@1。图中 Claude Opus 5 (xhigh) 以 67 分位居第一/并列第一,发帖者也直接评价 Anthropic 这次“做得很强”。
「模型」频道最新
- CogSci 讨论语言效率与 LLM 语义表征的关系 — ChrisGPotts · 2026-07-25
- Opus 5 无工具拿下 IMO 2026 42/42 — Afinetheorem · 2026-07-25
- Quadrillion 实测:Opus 5 在 ML 任务上快过 Opus 4.8 — igarciacamargo · 2026-07-25
- Google 在多数基准上落后于开源权重模型 — burny_tech · 2026-07-25
- 图表显示 Opus 5 在中等思考模式下编码性能最佳 — dejavucoder · 2026-07-25
- Opus 5、隐藏规则推断与 J-space 拼出新代理栈 — imjustnewatai · 2026-07-25