Agent Benchmarks Should Grade Like Resumes, Not One-Off Exams

A researcher argues that agent evaluation should shift from exam-style benchmarks that pose short problems like fixing a PR bug to resume-style assessments of long-horizon outcomes—such as autonomously deploying three B2C health apps to $700K ARR—and that post-training data should follow the same principle.

2026-08-16 ~ 2026-08-16 · 2 related posts