Sentry CEO: agent harness tests need real model calls, not just evals or mocks

zeeg · x · 2026-10-07

Sentry CEO David Cramer (zeeg) distinguishes evals (LLM-judged calls) from harness testing: building a harness means actually exercising the models you use in real inference, which he says is not the same thing as mock-based deterministic tests.

Related event: Sentry CEO: AI Agent Test Suite Costs $10 Per Run, Sparking Mock-vs-Real-Model Debate(8 posts)→

Original post →

More from coding & agent

coding & agent channel →