Nous Research launches Hermes Index agent leaderboard, Claude Opus 5.5 tops at 63.31
NousResearch · x · 2026-10-07
Nous Research introduced Hermes Index, an opinionated leaderboard averaging benchmarks to measure how models actually perform inside the Hermes Agent — meant to help users pick models and show labs their real agentic capability.
The index averages four suites: the new in-house Hermes Bench, Terminal-Bench 4.0, Terminal-Bench-Science, and SkillsBench. All models run the same harness with reasoning set to high where offered, reporting mean score and mean cost per task.
- Claude Opus 5.5 leads at 63.31, $4.99 per task
- GPT 6 Astra is second at 56.25 but priciest at $11.61 per task
- Claude Sonnet 5.5 follows at 53.14, $2.82
- At the low end: DeepSeek V4.1 Flash scores 36.91 at $0.26; Ling 3.0 Flash scores 21.56 at just $0.054
Related event: Nous Research Launches Hermes Index Agent Leaderboard, Claude Opus 5.5 Tops(4 posts)→
More from coding & agent
- 27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama — Ok-Shower7286 · 2026-10-07
- Google Testing Blog: Two-Way Doors — Don't Code Yourself into a Corner — rseroter · 2026-10-07
- Hot take: AGENTS.md and agent skills are just text files—put in whatever works — intellectronica · 2026-10-07
- Dev burns 190B tokens in a month using 30 AI subscriptions, $6k for $93k of API value — BLUECOW009 · 2026-10-07
- Set Up Your GrokBot Like Hiring a New Employee: One Job, Minimal Access — alex_verem · 2026-10-07
- ChatGPT Can Now Subscribe to Netlify Events: Auto-Reply Forms, Summarize Deploys — thisiskp_ · 2026-10-07