Evals shouldn't be skewed by infra: why turn limits break on sub-agents
langstonnashold · x · 2026-09-17
- @langstonnashold argues eval scores should never be influenced by infrastructure, but wall-clock time unfairly punishes bad infra: colocated sandboxes, fast/low-latency deployments, load spikes.
- Turn limits are also hard to apply to arbitrary harnesses: what counts as a turn when sub-agents are involved or the harness is an opaque binary?
- This is the other side of xeophon's proposal to use very long wall-clock caps (8h for Terminal-Bench 4.0).
More from coding & agent
- Monetize your MCP server: let agents pay via Stripe, already live at atomHQ and Donorbox — jeff_weinstein · 2026-09-17
- Supabase Select 26 lineup: YC's Garry Tan and Anthropic execs to speak in SF — garrytan · 2026-09-17
- Databricks rolled Astra out to all 3,500 engineers: beats Opus 5 on complex tasks, +60% coding spend — gdb · 2026-09-17
- Validating real-action agents is unsolved: one test run cost $30 and an X flag — Common_Dream9420 · 2026-09-17
- Texio: fail-closed Markdown section edits for coding agents, open-sourced under MIT — yjthegnius · 2026-09-17
- Agent Dev Pattern: Store a 'Consent Record' Before Every Agent Run — blaizedsouza · 2026-09-17