Self-modified harness matches Codex for $4.03, solving 82% of Terminal-Bench 2.1

omarsar0 · x · 2026-10-05

A method called SelfSearch has coding agents rewrite their own harness — instructions, tools, and procedures — using records of earlier self-modification attempts (reasoning, tool actions, outcomes), with each modified agent becoming the next improver.

Reported results: a self-modified harness matched Codex (top of a public nine-harness comparison) for just $4.03, solving 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search. Population-mean success rose across all six model-benchmark settings, with single agents gaining up to 11.2 points; on SWE-bench Multilingual one evolved agent gained 5.0 points.

Takeaway: optimizing your harness can squeeze far more performance than most teams realize.

Original post →

More from coding & agent

coding & agent channel →