Tokens and turns are useless eval metrics — only money and wall clock matter
xeophon · x · 2026-09-17
- xeophon's take: tokens and turns are artificial, hard to measure, and kinda useless metrics for agent evals; in the real world only two things matter — money spent and wall-clock time.
- His practice: set an absurdly high time limit as a backstop (8h for Terminal-Bench 4.0, 24h might be better) rather than capping turns/tokens.
More from coding & agent
- Routing simple requests to a cheap model made our total LLM bill worse — Massive_Tell_4276 · 2026-09-17
- Read-only analytics agent on Flash Lite: 2.2% cache hit rate, full cost data — RedaHaloubi · 2026-09-17
- Dev loves Codex app's /side chat so much he bound it to ⌘S — alex_frantic · 2026-09-17
- From personal second brain to team brain: one table, labels, and MCP access — Cole Medin · 2026-09-17
- AI SDK Adds evaluate Method, Welcomes Typesafe AI's New Jev Model — cramforce · 2026-09-17
- BotBell MCP lets AI assistants push notifications to iPhone and Mac — modelcontextprotocol · 2026-09-17