Atomic Agent outperforms Hermes on GAIA with the same Qwen-3.6-35B setup
kimmonismus · x · 2026-07-25
Atomic Agent beats Hermes on GAIA with the same Qwen model and Apple M4 Max
Atomic Agent reports stronger benchmark results than Hermes on GAIA Level 1 when both are run on the same 4-bit Qwen-3.6-35B model and the same Apple M4 Max machine.
- Atomic Agent: 37/53 tasks solved, 69.8%, finished in 3h 12m
- Hermes: 31/53 tasks solved, 58.5%, finished in 5h 10m
- Atomic solved 6 more tasks and finished nearly 2 hours faster
- The post frames Atomic as an open-source agent runtime under the MIT license
The benchmark details also suggest Atomic handled timeout-heavy tasks more efficiently than Hermes, spending less of its total runtime on failures.
Related event: Open-Source Atomic Agent Beats Hermes on GAIA Benchmark(3 posts)→
More from coding & agent
- Perplexity ships a CLI that gives coding agents web search access — AravSrinivas · 2026-07-25
- Opus 5 Codes 3D Colosseum Game with a Single Prompt — chrisfirst · 2026-07-25
- A simple proxy trick helps debug agent skills by intercepting every call — Daniel_Farinax · 2026-07-25
- GPT-5.6 Sol edges Opus 5 on DeepSWE with 72.7% vs 68.8% — rohanpaul_ai · 2026-07-25
- Nimbus launches as an open-source Astro framework for agent-ready docs — irvinebroque · 2026-07-25
- Creative Agency Workflow: Automating B-Roll Sourcing via Slack-Integrated AI Agent — beechinour · 2026-07-25