zeeg on $10 test suites: caching real model responses over mocks for harness testing

zwerp · x · 2026-10-07

David Cramer pushes back on mocked LLM responses for harness testing: many changes need real-model behavior, and separating the two adds huge complexity. He'll try a cached-inference approach first, responding to complaints that agent test suites cost $10 per run.

Related event: Sentry CEO: AI Agent Test Suite Costs $10 Per Run, Sparking Mock-vs-Real-Model Debate(8 posts)→

Original post →