Local open-source agent says it beats Hermes 37 to 31 on GAIA Level 1
kimmonismus · x · 2026-07-23
A local open-source agent claims to beat Hermes on benchmarks.
According to the post and leaderboard screenshot, the system:
- runs Qwen, Gemma, and Llama through llama.cpp
- uses stable-prefix caching to keep long sessions cheap
- applies TurboQuant to shrink the KV cache by 6.4×
- solved 37 tasks versus Hermes’ 31 on GAIA Level 1
The author argues this is evidence that local and open-source agents are becoming more competitive, including on macOS, Windows, and Linux.
More from coding & agent
- Claude is easier than the AWS console for building network topology HTML views — generativist · 2026-07-23
- Codex users report a frequent “thinking longer” delay over SSH — chaumian · 2026-07-23
- Kling shows an MCP workflow for cinematic AI video generation — Eric520CC · 2026-07-23
- GPT 5.6 Sol reportedly wrote a Laguna S2.1 inference implementation almost unaided — antirez · 2026-07-23
- Six practical ways to debug agent failures and keep performance from regressing — rdbms · 2026-07-23
- Carson Farmer says Codex can validate product hunches in hours — carsonfarmer · 2026-07-23