Evals turn 'good' into repeatable release checks for AI products

A tweet series argues that evals turn subjective notions of quality into repeatable release checks for AI products. Public results include Harvey nearly doubling internal quality scores after rebuilding its eval system, and Cursor cutting costs 41% via optimized model routing.

2026-09-23 ~ 2026-09-23 · 2 related posts