20 tasks × 3 repeats = 120 agent runs: the hidden cost of harness comparisons
RelationshipRound711 · reddit · 2026-10-02
Using Reef Infra's paired harness comparison as an example, the author tallies evaluation cost: documented episode count is 2 × tasks × episoderepeats — so 20 tasks and 3 repeats means 120 agent episodes, before counting the calls used to propose the edit. Each episode can itself contain multiple model and tool calls.
Key points:
- The two sides are the current harness and a candidate, same task suite, fixed model; repeats give multiple attempts per task and runs are interleaved.
- 120 is episodes executed, not a token budget or a statistical guarantee.
- As regression tasks accumulate, doubling the suite doubles this evaluation work; "no training GPU required" doesn't mean inference is free.
Advice: estimate cost from a few representative episodes before a long optimization run, then choose suite size and repeat count together — the proposal is only the visible tip of the workload.
More from coding & agent
- GPU still dominates agent app costs, not CPU sandboxes, says cost calculator — bookwormengr · 2026-10-02
- llama.cpp adds decision models: /v1/systemone scores options in a single forward pass — ggerganov · 2026-10-02
- 10 GitHub repos to level up your AI agents: from Browser Use to LangGraph and E2B — goyalshaliniuk · 2026-10-02
- Experimenting With Agentic AI to Create Vector Maps — rsasaki0109 · 2026-10-02
- Researcher: Agents that can work in a sandbox shouldn't have outbound access at all — moniquejmorrow · 2026-10-02
- Agents aren't GPU-bound: tool execution, memory bandwidth and sandbox overhead are the real bottleneck — ai · 2026-10-02