Harvey nearly doubled quality scores with evals; Cursor cut costs 41% via model routing
FinanceYF5 · x · 2026-09-23
Part 3 of an evals thread with public results: after reworking its eval system for its AI contract review tool, Harvey's internal quality scores nearly doubled. Cursor improved user satisfaction while cutting costs by 41% after optimizing its Auto Balance model routing.
Related event: Evals turn 'good' into repeatable release checks for AI products(2 posts)→
More from coding & agent
- Microsoft's Taste-Bench: best frontier model scores only 59.7% on long-horizon agent decisions — microsoft · 2026-09-23
- Lean Pool: an AI-agent-maintained archive of formalized mathematics — Vasily Ilin · 2026-09-23
- Indie dev builds an MCP server, hits 40-euro directory fee and blanket corporate IT blocks — sartomiki · 2026-09-23
- Model reviewers juggle 2-5 subscriptions per provider — corporate devs say that's not reality — emaayan · 2026-09-23
- Sergey Karayev: Google 'took itself out of the coding agent game' — sergeykarayev · 2026-09-23
- Running local agents in an isolated VM with open-webui and open-terminal — g1ccross · 2026-09-23