Sentry CEO: AI Agent Test Suite Costs $10 Per Run, Sparking Mock-vs-Real-Model Debate
Sentry CEO David Cramer (zeeg) took to Twitter on October 7 to complain: when an agent's test suite costs $10 in real LLM inference per run, what happens to "PR-maximization" style development that relies on AI agents filing high-frequency PRs. The complaint highlights a new problem as AI coding agents proliferate—agents hammering CI/tests drive inference costs up dramatically—and quickly sparked a debate over testing strategies for LLM applications.
Confirmed
- zeeg made clear the bulk of the bill comes from real inference costs, with the test suite costing roughly $10 per run.
- On testing methodology, zeeg distinguished two types of validation: evals are essentially calls that use an LLM to judge results, whereas testing an agent harness requires actually driving the model in use for real inference—the two are not the same thing.
- He insisted LLM calls should not be mocked, drawing an analogy to database calls—developers don't mock every database call; rather than mocking the "shape" of model outputs, it's better to run the real model, otherwise mocks break as soon as you swap models or behavior paths change.
- Facing cost pressure, zeeg said he plans to first try a compromise: caching inference responses—after a test passes, cache the model's responses and reuse them going forward instead of triggering real paid inference every time.
Not Confirmed
- Whether caching can significantly cut costs while preserving test validity remains unknown; zeeg only said he plans to "give it a try first," with no results yet.
Why It Matters
- QuavoBot, who joined the discussion, suggested caching model responses after the suite passes once and replaying them afterward—even caching for just a day could save substantial costs. Other developers argued that running real inference in harness tests introduces nondeterminism into otherwise deterministic tests, advocating mocks to cover edge cases. zeeg countered that many changes require observing their effect on real model behavior, and forcibly separating the two adds enormous complexity.
- The debate points to a common pain point in AI agent development: there is no industry consensus on the trade-off between testing fidelity and inference costs, and zeeg's caching compromise is currently one of the most-watched practical directions.
2026-10-07 ~ 2026-10-07 · 8 related posts
Primary sources
- [source] Your test suite costs $10 per run: the hidden CI bill of PR maxers — zeeg · 2026-10-07
- Stop paying real inference to run tests: cache model responses once the suite passes — QuavoBot · 2026-10-07
- Zeeg: when your test suite costs $10 per run, cache LLM responses instead of paying real inference — QuavoBot · 2026-10-07
- Skip Real Inference in Agent Tests: Mock LLM Responses for Determinism — zwerp · 2026-10-07
- What do PR maxxers do when your test suite costs $10 to run? — zeeg · 2026-10-07
- [source] zeeg on $10 test suites: caching real model responses over mocks for harness testing — zwerp · 2026-10-07
- [source] Sentry CEO: agent harness tests need real model calls, not just evals or mocks — zeeg · 2026-10-07
- Sentry CEO David Cramer: don't mock LLM calls, exercise real models in test harnesses — zeeg · 2026-10-07