Teknium pushes back on benchmark criticism: half of Hermes users use it for coding
Teknium · x · 2026-10-07
Teknium of Nous Research responded to criticism that Hermes' benchmark index overfits to terminal/science tasks irrelevant for a personal assistant. He clarified Hermes is not just an assistant, each capability has a separate leaderboard, and 50% of users use Hermes for coding. The exchange touches on how aggregated benchmark indices can mislead everyday users about model choice.
More from Models
- WebDev Arena: claude-opus-5.5-max tops leaderboard at 1814 with 846K votes cast — arena · 2026-10-07
- OpenAI releases frontier-model math results: Unique Games Conjecture and L=RL proved — burny_tech · 2026-10-07
- Mistral Large 4 Beats Opus 5.5 and GPT-6 Astra on Cybersecurity Benchmarks — Lower Refusal Rates — burny_tech · 2026-10-07
- Independent model Auro V9 ships, beating V8 head-to-head at a 10:1 clip — TheMoonMidas · 2026-10-07
- _xjdr: only four open-weights models are currently interesting — _xjdr · 2026-10-07
- Researcher _xjdr Praises Inkling Architecture, Calls DSV4 Models an Acquired Taste — _xjdr · 2026-10-07