HERMES: Executable Dev-Primitives Boost SWE Agent Performance by 12.4% at Lower Cost
Haibo Jin · hf · 2026-10-07
HERMES introduces Dev-Primitives, an abstraction that turns repository artifacts into active participants in software engineering, addressing brittleness of terminal-based agents on long-horizon workflows.
Design
- Each Dev-Primitive pairs a repo artifact with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication, and localized self-modification
- Dependency-aware dynamic activation instantiates primitives at repository scale
- A bug diagnosis mechanism maps execution evidence back to components that must be revised
Results
- Outperforms matched baseline harnesses by 12.4% on average across four SWE benchmarks
- Even with Qwen3-8B Dev-Primitives, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration while cutting inference cost by 26.2% on Terminal-Bench 4.0
- Conclusion: harness design matters as much as the underlying model
More from coding & agent
- Converting 10 coding harnesses including Claude Code into RL environments, cutting tool calls 31% — vllm_project · 2026-10-07
- AI-built websites are starting to look the same, developers blame shared design.md workflows — MickeySteamboat · 2026-10-07
- A mental model for AI assistants: dots own responsibilities, codex takes tasks — pvncher · 2026-10-07
- Soon people will ask 'How do you write code without a model?' — yunta_tsai · 2026-10-07
- Engineer shows what his screen looks like while AI agents do the work — KevinNaughtonJr · 2026-10-07
- Proofpress builds an evidence ledger for AI-agent research: withdrawing a finding revokes its dependents — tallmetommy · 2026-10-07