Open-Sourcing C&H: A Synthetic Law Firm to Benchmark Agent Continual Learning
mrdrozdov · x · 2026-08-10
Current agent benchmarks usually drop models into isolated task instances, failing to evaluate continual learning and memory. To address this, a team has open-sourced Calderwood & Harkness (C&H), a synthetic law firm environment.
Modeled after real legal work, C&H serves as a persistent environment where agents must perform multiple tasks over the same underlying context. It aims to test whether models can build on their experience—much like a senior engineer or lawyer leveraging a mental map and past playbooks to resolve issues effectively.
More from coding & agent
- OpenAI Caps Codex Context at 272k to Avoid High Cache-Read Costs — SquirrelMotor5379 · 2026-08-10
- AI Coding Agents Shift Developer Skills from Syntax to Architecture — ingliguori · 2026-08-10
- Why Do Agents Act Dumb? Dev Rants About Lack of Context Inference — ThatIsNotIllegal · 2026-08-10
- WolfBench: A Five-Metric Framework Redefining Agent Evaluation — morgymcg · 2026-08-10
- Open-Source Long Horizon Agent Harness with Cross-Session Memory & Sandboxing — Saboo_Shubham_ · 2026-08-10
- Conceptualizing a 'Gossip Protocol' for Human-Agent Communication — pzakin · 2026-08-10