Eval numbers shouldn't be skewed by infra: wall-clock time punishes bad setups

xeophon · x · 2026-09-17

In a debate on agent eval methodology, xeophon argues eval numbers should never be influenced by infrastructure, since wall-clock time overly punishes bad infra, colocated sandboxes, and load. He outlines what valid setups look like: an open model at low batch size on a Vera Rubin rack with sandboxes colocated, versus a 20-year-old server in Asia hitting Azure's US OpenAI endpoint (lowest TPS on OpenRouter).

Original post →

More from coding & agent

coding & agent channel →