Meta-Harness auto-searches harness code, gaining 4.7 points on IMO-level math
zainhas · x · 2026-09-14
A new arXiv paper, Meta-Harness (Chelsea Finn, Omar Khattab, et al.), introduces an outer-loop system that automatically optimizes the harness — the code deciding what info LLM apps store, retrieve, and present — using an agentic proposer that reads source code, scores, and execution traces of prior candidates via a filesystem.
Results: +7.7 points over a SOTA context-management system on online text classification with 4x fewer context tokens; a single discovered harness adds +4.7 points on 200 IMO-level problems averaged across five held-out models; discovered harnesses beat hand-engineered baselines on TerminalBench-2.
Poster zainhas also notes open-source agent framework goose is 'pretty good out of the box' and that harness selection hides lots of alpha.
More from coding & agent
- He spends $14k/month on AI subs and sells agent skills for $20/month — doodlestein · 2026-09-15
- C3 AI's DIA paper accepted at EMNLP 2026, tops all seven SQL benchmarks autonomously — C3_AI · 2026-09-15
- Give agents a goal: define 'done' and a budget, or 'keep trying' gets expensive — gethackteam · 2026-09-15
- GPT-6 vibecodes a full Silo fangame, silo scene added but gait still goofy — flngr · 2026-09-15
- OpenClaw's new 'suggested task' feature lets coding agents delegate work to themselves — steipete · 2026-09-15
- Dev Ships Tens of Millions of Lines of Memory-Safe Rust Agent Tooling on a 'Cancel-Correct' Runtime — doodlestein · 2026-09-15