Sentry Bets Big on Evals: 1,800 PR-Triggered Runs, 20-50% Token Cost Cuts
zeeg · x · 2026-09-23
- Sentry has invested heavily in an evals platform since May, driving a culture shift from shipping vibes to shipping with data.
- Past month stats: 1,800 eval runs triggered directly on PRs, 1,500 scenarios across 17 AI surfaces, 20-50% token cost reductions on several surfaces, and 10-20% eval pass-rate gains from core code-mode changes.
- Lenny Rachitsky's roundup: Ramp lifted receipt-collection accuracy from 35% to 83%; Shopify's AI workflow builder is 2.2x faster and 68% cheaper than its frontier-model setup; Harvey's rebuilt contract reviewer nearly doubled internal quality scores; Cursor cut Auto Balance routing costs 41% with higher satisfaction.
- Talent signal: nearly half of 25 PM job openings Lenny shared ask for eval-writing experience.
More from coding & agent
- Perplexity's hint-guided self-distillation cuts its agent's tool-call failures by 21.2% — perplexity_ai · 2026-09-23
- Tomo ships WebMCP Registry: open index of agent-ready tools across 100K sites — Jackyhuang · 2026-09-23
- LiteParse hits 2.8ms per PDF page, 25% faster in v2.14.6, claims fastest open-source parser — llama_index · 2026-09-23
- UCLA Releases ACLArena: A Framework for Agent Continual Learning in Multi-Stage Post-Training — UCLA-SCAI · 2026-09-23
- The Hidden Cost of Using Weaker Models for Decisions: You Never Benchmark, So You Never Notice the Lost Alpha — generativist · 2026-09-23
- Coinbase opens stock trading to AI agents, with x402 micropayments for live market data — MurrLincoln · 2026-09-23