Teknium pushes back on benchmark criticism: half of Hermes users use it for coding

Teknium · x · 2026-10-07

Teknium of Nous Research responded to criticism that Hermes' benchmark index overfits to terminal/science tasks irrelevant for a personal assistant. He clarified Hermes is not just an assistant, each capability has a separate leaderboard, and 50% of users use Hermes for coding. The exchange touches on how aggregated benchmark indices can mislead everyday users about model choice.

Original post →

More from Models

Models channel →