Harvey Open-Sources 100M+ Token Synthetic Law Firm Dataset for Agent Memory

marcbhargava · x · 2026-08-08

AI legal startup Harvey, in collaboration with EngramLab, has open-sourced a synthetic law firm dataset named Calderwood & Harkness, comprising over 100 million tokens of documents and 250 client cases.

While LLMs understand the law, they lack the proprietary experiential knowledge accumulated in actual firm practices. Furthermore, strict client data confidentiality prevents the naive use of real firm data for model training.

To solve this, the synthetic environment allows AI agents to handle multiple tasks while leveraging past case memory and experience. This enables better evaluation and training of agents' long-term memory and information retrieval in complex knowledge workflows.

Related event: Harvey Open-Sources 100M+ Token Synthetic Law Firm Dataset(3 posts)→

Original post →

More from coding & agent

coding & agent channel →