Meta researcher says coding-agent benchmarks are saturating and need goal-based evaluation

agihouse_org · x · 2026-07-28

Why coding-agent benchmarks need to move beyond task completion

Kilian Lieret, an AI research scientist at Meta and a contributor to SWE-agent, mini-SWE-agent, CodeClash, and ProgramBench, argues that agent benchmarks are reaching saturation if they only measure whether a task was completed.

The spotlight traces the evolution of evaluation from HumanEval (measurable code generation) to SWE-bench (real-world software engineering), and then explains why SWE-bench-style tests are no longer enough.

What should come next

Original post →

More from coding & agent

coding & agent channel →