Self-modified harness matches Codex for $4.03, solving 82% of Terminal-Bench 2.1
omarsar0 · x · 2026-10-05
A method called SelfSearch has coding agents rewrite their own harness — instructions, tools, and procedures — using records of earlier self-modification attempts (reasoning, tool actions, outcomes), with each modified agent becoming the next improver.
Reported results: a self-modified harness matched Codex (top of a public nine-harness comparison) for just $4.03, solving 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search. Population-mean success rose across all six model-benchmark settings, with single agents gaining up to 11.2 points; on SWE-bench Multilingual one evolved agent gained 5.0 points.
Takeaway: optimizing your harness can squeeze far more performance than most teams realize.
More from coding & agent
- Justine: "Simple, readable code" often just means missed edge cases — cnakazawa · 2026-10-06
- Google Docs now supports Markdown natively, turning files into shared agent memory — Saboo_Shubham_ · 2026-10-06
- 232x faster kernel with Codex auto-research: a GPU Mode contest postmortem — dejavucoder · 2026-10-06
- Claude Code picked Preact on its own — a glimpse of AI-driven tech stack decisions — tristanbob · 2026-10-06
- Group-Evolving Agents: a new paradigm where the unit of agent self-improvement is a group — xwang_lk · 2026-10-06
- "Nobody Is Vibe-Coding a Database" — Users Only Care If It Works — sujingshen · 2026-10-06