Paper: Re-Evaluate Production Agents with 38.5% of the Benchmark, Within 1.03 Points of Full Score
omarsar0 · x · 2026-09-24
A new paper shows how to re-evaluate a production agent at a fraction of the cost: 200 questions (38.5% of the full benchmark) reproduced the full score within 1.03 points. The study used 574 historical runs of an analytics agent serving tens of thousands of monthly users, comparing random sampling, caching, fixed subsets, and IRT-based adaptive testing. Multidimensional 2PL adaptive testing was most faithful, but the team deployed difficulty-stratified fixed subsets for simplicity—they transferred to five other agent families without recalibration and stayed stable with calibration windows as short as one day.
More from coding & agent
- Resend joins Stripe Projects: one CLI command wires email into AI agent workflows — jeff_weinstein · 2026-09-24
- MiniMax Code tested: building and deploying a full site from a single prompt — nikola_mr64990 · 2026-09-24
- Agent Reliability QA: The Next AI Category Is Stress-Testing Agents, Not Building Them — DevWithTea123 · 2026-09-24
- Privacy Middleware Concept: Anonymize Sensitive Data Locally Before It Reaches LLMs — Ai_MOON_SHOT · 2026-09-24
- No-BS guide to agent harness engineering: context, guardrails, and feedback — Pavan_Belagatti · 2026-09-24
- Claude Code Projects adds local support: threads can now run on your own machine — bcherny · 2026-09-24