Harness Arena: open-source blind benchmark for AI coding agent harnesses
Due_Armadillo_8744 · reddit · 2026-09-02
A Reddit developer open-sourced Harness Arena (MIT license) to fix agent harness comparisons that change too many variables at once — different model, prompt, tools, environment and task.
- Design: harnesses get the same task, run independently in isolated workspaces, deliverables are collected, outputs shown anonymously, scored before identities are revealed, and results feed an Elo leaderboard
- Integrating/testing Claude Code, Codex CLI, Hermes, OpenClaw, OpenCode and OnDemand, with free Kimi, Qwen and GLM options for experiments
- Open methodology questions: same model? same reasoning effort? equal token/context budget, identical tools/MCP servers, wall-clock/retry/cost limits? Should latency and cost be scored or shown separately? Category-specific leaderboards vs one overall board?
- Repo: github.com/Ondemand-OSS/harness-arena; the author is soliciting input on what methodology would make a harness-vs-harness benchmark trustworthy
Related event: Harness Arena: Open-Source Blind Benchmark for Coding Agents(2 posts)→
More from coding & agent
- Unitpost launches on Product Hunt: one email platform AI agents can send through via MCP — thisiskp_ · 2026-09-02
- Google ships agentic video understanding with Gemini, explainer video released — patloeber · 2026-09-02
- Post-training, custom spec decoding and vLLM tuning: a hands-on inference cost-saving playbook — dhruv2038 · 2026-09-02
- Simon Willison vibe-codes a GeoJSON map tool, with ChatGPT Work pulling government boundary data — MaxLenormand · 2026-09-02
- Dev on the Codex hype: less unreliable than claimed, with marketing on both sides — D3VAUX · 2026-09-02
- Developer runs a 'reality-check' Claude Code skill on every new frontier model to audit his half-finished projects — doodlestein · 2026-09-02