Frontier-Bench 图表对比 Claude Opus 5、Fable 5 和 GPT-5.6 Sol 编程表现

thesaraharminta · x · 2026-07-25

配图是一张 Frontier-Bench v0.1 的 agentic coding 图,按不同 effort level 对比了 Claude Opus 5、Claude Fable 5、Claude Opus 4.8、GPT-5.6 Sol。

这张图想表达的是:Opus 5 在提高 effort 后表现继续上升,最高大致到低 40% 区间;与此同时,Fable 5 和 GPT-5.6 Sol 也体现出不同的成本—表现权衡。

所属事件:Anthropic 发布 Claude Opus 5,多项基准测试领先(34 条相关)→

原文链接 →

「模型」频道最新

更多「模型」频道 AI 资讯 →