FrontierHarness eval: 12 agent harnesses, same model — pass rates span 50%–67%, cost per pass $1.05–$18.34
stuffyokodraws · x · 2026-09-03
- The FrontierHarness eval argues the key question has shifted from "which model" to "which harness."
- Setup: 12 harnesses (Pi, Exo, Claude Code, Codex, DeepSeek Harness and more) tested on the same model, tasks and runtime — 360 runs, 2 billion tokens.
- Results: pass rates ranged from 50% to 67%, while cost per pass varied from $1.05 to $18.34 — harness engineering now matters as much as model choice for both quality and cost.
More from coding & agent
- Omnara: open-source, self-hostable alternative to Claude managed agents — JaynitMakwana · 2026-09-03
- NanoCodana: open-source Claude Code-style coding agent that runs entirely in the browser — andrepimentaa7 · 2026-09-03
- Solo user builds governance-first multi-agent system: lead agent, least privilege, independent auditor — Grimmoner · 2026-09-03
- Deploy the Foundry Model Router with Azure Bicep — adnan_hashmi · 2026-09-03
- Code Arena launches WebDev Pareto frontier to rank AI models by quality vs price — arena · 2026-09-03
- reverse-skill: 30k-Star GitHub repo routes AI agents through 44 reverse-engineering skill playbooks — alex_verem · 2026-09-03