Two lawyers agreed on only 45% of 500 contract edits — how Crosby builds legal AI evals
graceisford · x · 2026-10-08
Aveek Duttagupta, founding engineer at AI law firm Crosby, details how the company builds evals for legal reasoning.
Context
- Crosby runs its own law firm, aiming for agents to do as much legal work as possible, delivering in hours what takes weeks; lawyers and engineers build agents side by side.
- Contracts are the entry point: high volume, cross-industry, grounded in concrete context (the contract, client positions, negotiation history, commercial risk) — but the nuance of weighing them exceeds the documents alone.
The core puzzle: subjective yet measurable
- In one experiment, two skilled attorneys each created ideal redlines (golden set) for the same 20 NDAs; of 500 total edits they agreed on only 45%.
- Unlike traditional software, no compiler or unit tests can tell whether an AI review is something a human would send to a client — turning this non-verifiable problem into something verifiable at scale is their core engineering challenge.
A concrete case study of agent eval engineering in a non-verifiable domain.
Related event: Legal AI Firm Crosby Explains Building Evals for Legal Reasoning(2 posts)→
More from coding & agent
- Tavily, LangChain and Nebius host SF Tech Week event with 100K credits for startups — LangChain · 2026-10-08
- Dev ditches OpenClaw for Grok bot, citing speed and free live X access — haltakov · 2026-10-08
- Obsidian Starter Kit v5 automates archiving, todos and AI context so you stop doing them by hand — dSebastien · 2026-10-08
- Hamel Husain: same model as judge is usually fine — verify human-label alignment first — HamelHusain · 2026-10-08
- Mitchell Hashimoto's Rex: a terminal replacement built for AI coding agents — ricklamers · 2026-10-08
- DAIR.AI curates 31 papers on recursive self-improvement, from Gödel Machines to today — omarsar0 · 2026-10-08