Agent Arena 榜单:同档智能体任务成本相差超五倍
arena · x · 2026-08-20
Agent Arena 发布按「净提升 × 每任务成本」的智能体模型性价比榜单(基于真实 Agent 任务,187 万会话、49 个模型)。
- Pareto 最优:Claude Opus 5 (High) $1.78/任务、+12.34% 排第一;Kimi K3 (Max) 以 $0.62/任务、+10.53% 位列性价比黑马;GPT 5.5 (High) $0.44/任务、+7.75%;Grok 4.5 $0.22/任务、+6.08%;GLM 5.2 (Max) $0.18/任务、+6.06%。
- 前五名内每任务中位成本从 Kimi K3 的 $0.62 到 Claude Opus 5 (Max) 的 $3.37,相差超过 5 倍。
- Claude Opus 5 (Max) 比 Opus 5 (High) 贵近一倍,得分却略低(+11.97% vs +12.34%)。
所属事件:Agent Arena 榜单:同档智能体任务成本相差超五倍(2 条相关)→
「模型」频道最新
- Qwen3.8-27B 在 AMD Strix Halo 上跑出 31t/s — stereohype · 2026-08-20
- incoai 推出 Qwen3.8-27B-DFlash2 草稿模型,主打投机解码加速 — incoai · 2026-08-20
- 实测:Ornith-1.5 9B 低配环境表现优异 — zippydazoop · 2026-08-20
- 本地跑 Qwen 27B 写出 3D 房间场景,网友惊呼时代变了 — dsdt · 2026-08-20
- Superwhisper 推出 0.6B 端侧模型 S1-mini — Scobleizer · 2026-08-20
- Agent Arena 榜单:Kimi K3 性价比第一,Claude Opus 5 最贵 — arena · 2026-08-20