Hermes Index scores: Opus 5.5 at $4.99/task leads, GPT 6 Astra costs $11.61
NousResearch · x · 2026-10-07
Nous Research shared full Hermes Index results — same harness for every model, reasoning maxed where offered, reporting mean score and mean cost per task across four suites.
- Claude Opus 5.5: 63.31 at $4.99/task
- GPT 6 Astra: 56.25 at $11.61/task (most expensive by far)
- Claude Sonnet 5.5: 53.14 at $2.82/task
- DeepSeek V4.1 Flash: 36.91 at $0.26/task
- Ling 3.0 Flash: 21.56 at $0.054/task
The spread tells a clear story: Sonnet 5.5 delivers 53 points at under 60% of Opus's cost, while budget Flash models trade near-zero cost for much lower scores.
Related event: Nous Research Launches Hermes Index Agent Leaderboard, Claude Opus 5.5 Tops(4 posts)→
More from coding & agent
- 27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama — Ok-Shower7286 · 2026-10-07
- Google Testing Blog: Two-Way Doors — Don't Code Yourself into a Corner — rseroter · 2026-10-07
- Hot take: AGENTS.md and agent skills are just text files—put in whatever works — intellectronica · 2026-10-07
- Dev burns 190B tokens in a month using 30 AI subscriptions, $6k for $93k of API value — BLUECOW009 · 2026-10-07
- Set Up Your GrokBot Like Hiring a New Employee: One Job, Minimal Access — alex_verem · 2026-10-07
- ChatGPT Can Now Subscribe to Netlify Events: Auto-Reply Forms, Summarize Deploys — thisiskp_ · 2026-10-07