Harvey Open-Sources 100M+ Token Synthetic Law Firm Dataset for Agent Memory
marcbhargava · x · 2026-08-08
AI legal startup Harvey, in collaboration with EngramLab, has open-sourced a synthetic law firm dataset named Calderwood & Harkness, comprising over 100 million tokens of documents and 250 client cases.
While LLMs understand the law, they lack the proprietary experiential knowledge accumulated in actual firm practices. Furthermore, strict client data confidentiality prevents the naive use of real firm data for model training.
To solve this, the synthetic environment allows AI agents to handle multiple tasks while leveraging past case memory and experience. This enables better evaluation and training of agents' long-term memory and information retrieval in complex knowledge workflows.
Related event: Harvey Open-Sources 100M+ Token Synthetic Law Firm Dataset(3 posts)→
More from coding & agent
- DeepSeek Cascade Beats GPT-5.6 Luna on DeepSWE at 37% Lower Cost — togethercompute · 2026-08-08
- Guardrails Hinder Defense: Dev Calls for Open Models to Harden Security — max_paperclips · 2026-08-08
- Can MiniMax H3 Handle Automated Video Generation Pipelines? — datavyro · 2026-08-08
- Open-Source Agent 'sol-advisor' Upgrades with Support for Cursor, Copilot and More — daniel_mac8 · 2026-08-08
- Parahelp Launches Automation: Auto-Sync AI Support Agents with Codebase Updates — Scobleizer · 2026-08-08
- CopilotKit Launches OpenTag: Open-Source Agent Triage for Slack & Teams — Roger_M_Taylor · 2026-08-08