Frontier-Bench 图表对比 Claude Opus 5、Fable 5 和 GPT-5.6 Sol 编程表现
thesaraharminta · x · 2026-07-25
配图是一张 Frontier-Bench v0.1 的 agentic coding 图,按不同 effort level 对比了 Claude Opus 5、Claude Fable 5、Claude Opus 4.8、GPT-5.6 Sol。
这张图想表达的是:Opus 5 在提高 effort 后表现继续上升,最高大致到低 40% 区间;与此同时,Fable 5 和 GPT-5.6 Sol 也体现出不同的成本—表现权衡。
所属事件:Anthropic 发布 Claude Opus 5,多项基准测试领先(34 条相关)→
「模型」频道最新
- 新前沿图表显示 Grok 4.5、SWE-1.7 与 Opus 5 同场竞争 — GavinSBaker · 2026-07-25
- Anthropic 发布 Claude Opus 5,定价不变但编码分数更强 — kimmonismus · 2026-07-25
- Opus 5 评测显示,5-agent 编程团队到 0.6 分快 2.2 倍 — OfirPress · 2026-07-25
- 用户称 Fable 5 仍胜过被追捧的 Claude 5 — MicahBerkley · 2026-07-25
- Claude Opus 5 与 5 Fast 上线,主打长任务执行 — matanSF · 2026-07-25
- Claude Opus 5 据称性能接近 Fable,价格只有一半 — Steap-Edit · 2026-07-25