Eval numbers shouldn't be skewed by infra: wall-clock time punishes bad setups
xeophon · x · 2026-09-17
In a debate on agent eval methodology, xeophon argues eval numbers should never be influenced by infrastructure, since wall-clock time overly punishes bad infra, colocated sandboxes, and load. He outlines what valid setups look like: an open model at low batch size on a Vera Rubin rack with sandboxes colocated, versus a 20-year-old server in Asia hitting Azure's US OpenAI endpoint (lowest TPS on OpenRouter).
More from coding & agent
- Databricks rolled Astra out to all 3,500 engineers: beats Opus 5 on complex tasks, +60% coding spend — gdb · 2026-09-17
- Validating real-action agents is unsolved: one test run cost $30 and an X flag — Common_Dream9420 · 2026-09-17
- Texio: fail-closed Markdown section edits for coding agents, open-sourced under MIT — yjthegnius · 2026-09-17
- Texio: fail-closed Markdown section edits for coding agents, open-sourced under MIT — yjthegnius · 2026-09-17
- Agent Dev Pattern: Store a 'Consent Record' Before Every Agent Run — blaizedsouza · 2026-09-17
- A practical prompt recipe for AI code review: solve it yourself first, then compare — dotey · 2026-09-17