New “expenditure horizon” metric compares human and agent cost efficiency on open-ended tasks
littmath · x · 2026-07-22
A proposed way to measure AI agent capability on continuously scored tasks: the expenditure horizon.
- It compares how human and agent performance improve as budget increases on open-ended optimization problems.
- The key point is the spend level where a human becomes more cost-effective than the agent.
- That crossover budget is defined as the agent’s expenditure horizon.
Related event: METR Proposes 'Expenditure Horizon' for AI Agent Evaluation(2 posts)→
More from Research
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11
- Negative Self-Distillation improves LLM reasoning by avoiding flawed reasoning paths — Rongcan Pei · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- Yann LeCun live at ECCV on World Models — Weak_Assistance_5261 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11