Claude Opus 5 在编码、搜索和推理基准上全面走强
natolambert · x · 2026-07-25
Anthropic 的 Claude Opus 5 被称为在“更快迭代速度 + 规模化 RL”下表现显著提升的前沿模型。
配图给出了多项基准对比,Opus 5 在多个任务上领先或接近领先,包括:
- Agentic terminal coding:43.3%
- GDPVal-AA v2:1861
- ARC-AGI-3:30.2%
- BrowseComp:90.8%
- Humanity’s Last Exam:无工具 56.3%,有工具 64.7%
- OSWorld 2.0:70.6%
- FrontierCode v1.1 Main:53.4%
- AutomationBench:26.0%
- Legal Agent Benchmark:11.7%
- HealthBench Professional:59.8%
- BioMysteryBench:hard 49.4%,human solved 90.1%
同时,帖子提到其安全分类器的介入频率预计会比 Fable 5 低约 85%。
所属事件:Anthropic 发布 Claude Opus 5(35 条相关)→
「模型」频道最新
- Anthropic 发布 Claude Opus 5,主打生物和化学能力 — gabepgomes · 2026-07-25
- Claude Opus 5 连自家系统卡都拒绝引用,版权边界很严 — peterwildeford · 2026-07-25
- Nvidia 将在直播中解析 Nemotron 3 Ultra 开源模型 — yacinelearning · 2026-07-25
- Claude Opus 5 在 LiveBench 逼近前排,实测仍落后 Fable 5 — bindureddy · 2026-07-25
- 有人调侃 Opus 5 只比 Fable 5 强一点,但价格砍半 — dejavucoder · 2026-07-25
- Every 对 Opus 5 转向负评,上月还在夸 Opus 4.8 — soumitrashukla9 · 2026-07-25