Nous Research 发布 Hermes Index 智能体榜单,Claude Opus 5.5 居首

Nous Research 发布 Hermes Index 榜单,衡量各模型在 Hermes Agent 中的实际表现,帮助用户选模型,也让实验室了解自身定位。榜单对四套基准取平均,其中包括新自研的 Hermes Bench,其余为 Terminal-Bench 等。测试统一使用同一套 harness,凡有推理档位的模型一律开到最高强度,并报告单任务成本:头部模型虽贵但表现领先,尾部模型单任务成本可低至 5 美分。Claude Opus 5.5 位列榜首。

2026-10-07 ~ 2026-10-07 · 4 条相关