Adobe paper: grading agents against live code-verified answers cuts errors 29%
rohanpaul_ai · x · 2026-10-01
A new Adobe paper tackles a flaw in standard agent evaluations: they assume the correct answer never changes, which breaks on live data science tasks where ground truth shifts over time.
The fix: express the expected answer as executable code that fetches the current correct result each time the test runs, then have an AI grader check the agent's reply against that code output.
Results on 53 test cases: the AI grader matched human experts 29% better with code-based references than with written descriptions, while using 16% fewer tokens. With no reference answer at all, the grader performed worse than random.
Paper: "Skill-based Agentic Evaluation for Real-time Data Science Tasks" (arXiv:2609.16487).
More from Research
- NormViz benchmark: best model Gemini 3 Flash scores just 25.3% on visual cultural norms across 16 countries — StellaLisy · 2026-10-01
- Workspace Models: Memory Architecture Turns Reasoning Agents into Robot Policies — jajoosam · 2026-10-01
- Dunning-Kruger's original paper is "really bad" research, argues dev backed by Sabine Hossenfelder — dbasch · 2026-10-01
- Princeton professor: AI safety has long been dismissed despite serious open questions — HazanPrinceton · 2026-10-01
- DevDataLab launches open urban data platform for 10,000 cities, aiming to be the Penn World Table for cities — paulnovosad · 2026-10-01
- SWE-sweep benchmark: top models score under 5% finding bugs unsupervised — OfirPress · 2026-10-01