METR defines an “expenditure horizon” for AI optimization, using NanoGPT as a test case
gleech · x · 2026-08-04
METR proposes an “expenditure horizon” metric for AI optimization ability
METR introduces a way to measure how cost-effective AI agents are at optimization tasks by comparing human and agent performance as a function of total expenditure, including token cost, compute, and human labor.
- They define the “expenditure horizon” as the budget point where human effort becomes more cost-effective than AI.
- As an empirical example, they apply the method to the NanoGPT speedrun.
- Their estimate suggests that, on the margin, each 1% improvement in NanoGPT costs about $2,500 in human labor.
- For preliminary agentic optimization runs, METR says that after more than $10K of expenditure, the estimated expenditure horizon is around $0–$3K.
The broader motivation is whether AI is accelerating AI R&D, and by how much. The authors argue this is hard to measure directly, so they use this cost-based framework instead.
More from Research
- A model-stealing paper’s proof needs a stronger spanning assumption, the author says — ArthurConmy · 2026-08-04
- Model-stealing paper author says the original theorem was wrong, but the result can be repaired — ArthurConmy · 2026-08-04
- RedMonk says only 1% of surveyed open source code is machine-written — rseroter · 2026-08-04
- Robot walking demo adds target-facing rewards but still fails 80% of the time — carlosdponx · 2026-08-04
- Airbnb details its internal AI eval stack, sampling 5% of traffic daily — econoar · 2026-08-04
- Aero Hand Open debuts as a $314 open-source robotic hand with 16 joints — TinfoilTricorn · 2026-08-04