Sentry founder on eval costs: keep asserts deterministic, LLM judges get brutal
zeeg · x · 2026-09-10
In an exchange with another developer, Sentry founder David Cramer (zeeg) shares how he keeps LLM eval costs down: most of his assertions are deterministic, rubrics are minimal, and only a few non-deterministic asserts require running an agent. The counterpart complains that testing 10 models across 3 evals with LLM-as-judge is already expensive, and worries costs will balloon 100x with more evals and pricier models — a candid look at how brutal agent eval iteration costs have become.
Related event: Sentry Founder Shares LLM Eval Cost Lessons(3 posts)→
More from coding & agent
- Running 3 parallel vibecad instances tripled one engineer's design iteration speed — IanPritchard · 2026-09-10
- 2026 Survey of Agents opens for responses to map the evolving agent ecosystem — AlexGDimakis · 2026-09-10
- You can just ask ChatGPT to find spherical codes — felpix_ · 2026-09-10
- Harness optimization lifts Harvey legal agent benchmark pass rate from 67.1% to 85.9% — sarahookr · 2026-09-10
- Dev recreates 1X hand demo from scratch with Astra agent in MuJoCo simulation — TheMoonMidas · 2026-09-10
- Six-year LLM-assisted Melee decompilation powers stunning MR demo on Meta Quest — TheMoonMidas · 2026-09-10