Sentry CEO: agent harness tests need real model calls, not just evals or mocks
zeeg · x · 2026-10-07
Sentry CEO David Cramer (zeeg) distinguishes evals (LLM-judged calls) from harness testing: building a harness means actually exercising the models you use in real inference, which he says is not the same thing as mock-based deterministic tests.
More from coding & agent
- Dev slams Linear for shipping "just a chat box" instead of rethinking work for AI agents — devenbhooshan · 2026-10-07
- Dev claims 700,000 lines of code in 4 days — "just me and Opus" — haydendevs · 2026-10-07
- Opus 5.5 nearly matches Fable 5.1 at two-thirds the cost in real-repo bug benchmark — PawelHuryn · 2026-10-07
- Bug Hunt Benchmark yields different rankings for frontier models — PawelHuryn · 2026-10-07
- DAEDALUS bootstraps agent memory from self-generated tasks, +15.9 points success rate — illuin · 2026-10-07
- Indie dev finds Claude Code driving Figma beats vibe coding for SEO landing pages — yihui_indie · 2026-10-07