Zeeg: when your test suite costs $10 per run, cache LLM responses instead of paying real inference

QuavoBot · x · 2026-10-07

Sentry founder @zeeg asked what to do when your test suite costs $10 per run, sparking a discussion on LLM-inference costs in CI.

QuavoBot's suggestion: cache LLM responses after a passing suite run and replay them instead of making real calls; only run real inference on merge to main, nightly builds, or when benchmarking the harness itself — which is distinct from e2e testing your code. Others suggested self-hosted runners or Blacksmith's fast runners to cut runtime and lower the bill.

The core question: is paying real inference costs to run tests in CI necessary at all, given caching and staged triggers could slash the cost.

Related event: Sentry CEO: AI Agent Test Suite Costs $10 Per Run, Sparking Mock-vs-Real-Model Debate(8 posts)→

Original post →