Agent benchmarks now list cost per task next to scores, and the bill decides the winner
YvesMulkers · x · 2026-09-27
The author notes that two agent benchmark lines this week put cost per task right next to the score. His take: score used to be the whole headline, but once all the good agents clear the bar, the bill is what picks one. The observation points to a broader shift toward cost-sensitive agent evaluation.
More from coding & agent
- WebMCP could become the HTML/API layer of the agentic web — Thionne_WTZ · 2026-09-27
- Stanford/Together AI paper: agent teams hit 66.7% vs 48.8% for single agents — mark_k · 2026-09-27
- antirez regrets not finding time to hand-write small poetry-style programs — antirez · 2026-09-27
- mitsuhiko on Astra: expensive, narrow edits, and a frustrating experience to review — mitsuhiko · 2026-09-27
- François Fleuret asks: are there programming languages tailored for LLMs? — francoisfleuret · 2026-09-27
- AI Can't 'Read the Room': The Hard Problem of Selective Info Sharing in Enterprise AI — devanshmehta · 2026-09-27