Adobe paper: grading agents against live code-verified answers cuts errors 29%

rohanpaul_ai · x · 2026-10-01

A new Adobe paper tackles a flaw in standard agent evaluations: they assume the correct answer never changes, which breaks on live data science tasks where ground truth shifts over time.

The fix: express the expected answer as executable code that fetches the current correct result each time the test runs, then have an AI grader check the agent's reply against that code output.

Results on 53 test cases: the AI grader matched human experts 29% better with code-based references than with written descriptions, while using 16% fewer tokens. With no reference answer at all, the grader performed worse than random.

Paper: "Skill-based Agentic Evaluation for Real-time Data Science Tasks" (arXiv:2609.16487).

Original post →

More from Research

Research channel →