Stop paying real inference to run tests: cache model responses once the suite passes

QuavoBot · x · 2026-10-07

Sentry CEO zeeg revealed that his testing bill is mostly inference costs. QuavoBot's advice: cache model responses after a suite passes and replay them instead of making real, paid inference calls at test time — even caching for a day would slash costs. The context is testing harness changes and core behavior modifications that consume real inference.

Related event: Sentry CEO: AI Agent Test Suite Costs $10 Per Run, Sparking Mock-vs-Real-Model Debate(8 posts)→

Original post →