Researchers use Item Response Theory to predict coding agent performance on unseen tasks
dan_fried · x · 2026-10-09
Presented at COLM 2026, this work applies Item Response Theory (IRT) to coding agent evaluation:
- Today's agent evals reduce everything to a single benchmark accuracy, obscuring which tasks are harder and why
- The study analyzes agent performance at the task level and predicts how new agents perform on new tasks, including unseen model/harness combos
- This reduces eval noise from harness swaps and makes leaderboard results more transferable
The paper will appear at the ICLR 2026 AIWILD workshop.
More from coding & agent
- Hone launches Engines: AI that owns business outcomes, with simulation for self-improving harnesses — graceisford · 2026-10-09
- LangChain Podcast Digs Into Decision Models as OpenAI and Databricks Enter the Space — LangChain · 2026-10-09
- Persistent Self-Improvement: Why Having Memory Isn't the Same as Learning — anirudhg9119 · 2026-10-09
- Ruff Author Confirms His Tools Are Tested Almost Entirely via Black-Box Suites — charliermarsh · 2026-10-09
- How to Test New AI Models on Your Own Tasks, Weighing Quality, Speed and Cost — The AI Daily Brief · 2026-10-09
- Leak hints OpenAI runs multi-agent swarms on time budgets, not token budgets, and treats it as IP — maksym_andr · 2026-10-09