Agent benchmarks should work like resumes: open-ended org goals, not one-shot PR fixes

menhguin · x · 2026-08-16

A researcher contrasts two paradigms for evaluating agents. Exam-style benchmarks scope precisely which specific short-horizon solutions to pursue — e.g. "solve this PR bug scenario" — which only tests one-shotting a slice of a codebase. A "resume"-style evaluation, by contrast, measures whether an agent can agentically match complex organizational-level goals against a conventional but highly open-ended rubric.

Concretely, he proposes tasks like "autonomously deploy 3 B2C health applications with $700k ARR" — closer to a resume, and gradable as such. He argues the latter kind of evaluation is valuable right now.

Related event: Agent Benchmarks Should Grade Like Resumes, Not One-Off Exams(2 posts)→

Original post →

More from coding & agent

coding & agent channel →