Agent eval harnesses: why one researcher favors 8h wall-clock limits over turn caps
xeophon · x · 2026-09-17
- xeophon argues agent evals shouldn't impose artificial turn/token limits, instead settling on very high wall-clock time caps: 8 hours for Terminal-Bench 4.0, with 24 hours possibly better.
- Counterpoint from @langstonnashold: wall clock over-punishes infra variance (colocated sandboxes, low-latency deployments, load), while turn counts are hard to define once sub-agents or opaque harness binaries are involved.
More from coding & agent
- Monetize your MCP server: let agents pay via Stripe, already live at atomHQ and Donorbox — jeff_weinstein · 2026-09-17
- Supabase Select 26 lineup: YC's Garry Tan and Anthropic execs to speak in SF — garrytan · 2026-09-17
- Databricks rolled Astra out to all 3,500 engineers: beats Opus 5 on complex tasks, +60% coding spend — gdb · 2026-09-17
- Validating real-action agents is unsolved: one test run cost $30 and an X flag — Common_Dream9420 · 2026-09-17
- Agent Dev Pattern: Store a 'Consent Record' Before Every Agent Run — blaizedsouza · 2026-09-17
- A practical prompt recipe for AI code review: solve it yourself first, then compare — dotey · 2026-09-17