20 tasks × 3 repeats = 120 agent runs: the hidden cost of harness comparisons

RelationshipRound711 · reddit · 2026-10-02

Using Reef Infra's paired harness comparison as an example, the author tallies evaluation cost: documented episode count is 2 × tasks × episoderepeats — so 20 tasks and 3 repeats means 120 agent episodes, before counting the calls used to propose the edit. Each episode can itself contain multiple model and tool calls.

Key points:

Advice: estimate cost from a few representative episodes before a long optimization run, then choose suite size and repeat count together — the proposal is only the visible tip of the workload.

Original post →

More from coding & agent

coding & agent channel →