High-signal evals beat harness tweaks beat post-training: a layered agent methodology
abeirami · x · 2026-09-10
A discussion thread lays out a layered optimization framework for agent engineering: evals > harness optimization (fast learning) > model post-training (slow learning), arguing most agentic tasks never need slow learning loops. Key points: define success first and build high-SNR evals to map your agent's current frontier; use evals to attribute how much intelligence a task needs and build model routing policies to save cost; even when post-training, first get high-SNR evals, optimize the harness on that signal, then "distill" the system behavior back into weights.
Related event: Researcher Proposes Agent Capability Pyramid: Evals Over Post-Training(2 posts)→
More from coding & agent
- FrogNano: a 4B model trained purely with RL on synthetic tasks hits repo-level coding — burkov · 2026-09-10
- Instinct launches Trusted Person network letting AI agents negotiate plans with each other — mon__lim · 2026-09-10
- diagram-design skill gives Codex and Claude Code 39 actually-good diagram types — daniel_mac8 · 2026-09-10
- Simular to host CUA party at SF Tech Week debating if 2027 brings AGI for computer-use agents — xwang_lk · 2026-09-10
- LangChain's Harrison Chase: code-writing LLMs will make computer use take off — hwchase17 · 2026-09-10
- Open-source Chrome extension lets a browser agent keep working in its tab while you browse elsewhere — Silly_Entertainer92 · 2026-09-10