Meta-Harness auto-searches harness code, gaining 4.7 points on IMO-level math

zainhas · x · 2026-09-14

A new arXiv paper, Meta-Harness (Chelsea Finn, Omar Khattab, et al.), introduces an outer-loop system that automatically optimizes the harness — the code deciding what info LLM apps store, retrieve, and present — using an agentic proposer that reads source code, scores, and execution traces of prior candidates via a filesystem.

Results: +7.7 points over a SOTA context-management system on online text classification with 4x fewer context tokens; a single discovered harness adds +4.7 points on 200 IMO-level problems averaged across five held-out models; discovered harnesses beat hand-engineered baselines on TerminalBench-2.

Poster zainhas also notes open-source agent framework goose is 'pretty good out of the box' and that harness selection hides lots of alpha.

Related event: Stanford/MIT Paper: Same Model Can Perform 6x Worse With a Different Harness(3 posts)→

Original post →

More from coding & agent

coding & agent channel →