Paper: Re-Evaluate Production Agents with 38.5% of the Benchmark, Within 1.03 Points of Full Score

omarsar0 · x · 2026-09-24

A new paper shows how to re-evaluate a production agent at a fraction of the cost: 200 questions (38.5% of the full benchmark) reproduced the full score within 1.03 points. The study used 574 historical runs of an analytics agent serving tens of thousands of monthly users, comparing random sampling, caching, fixed subsets, and IRT-based adaptive testing. Multidimensional 2PL adaptive testing was most faithful, but the team deployed difficulty-stratified fixed subsets for simplicity—they transferred to five other agent families without recalibration and stayed stable with calibration windows as short as one day.

Original post →

More from coding & agent

coding & agent channel →