Nous Research launches Hermes Index agent leaderboard, Claude Opus 5.5 tops at 63.31

NousResearch · x · 2026-10-07

Nous Research introduced Hermes Index, an opinionated leaderboard averaging benchmarks to measure how models actually perform inside the Hermes Agent — meant to help users pick models and show labs their real agentic capability.

The index averages four suites: the new in-house Hermes Bench, Terminal-Bench 4.0, Terminal-Bench-Science, and SkillsBench. All models run the same harness with reasoning set to high where offered, reporting mean score and mean cost per task.

Related event: Nous Research Launches Hermes Index Agent Leaderboard, Claude Opus 5.5 Tops(4 posts)→

Original post →

More from coding & agent

coding & agent channel →