Baseten and Harvey boost M&A diligence agent pass rate from 29% to 63% with RL on RLM harness
baseten · x · 2026-09-09
Baseten partnered with legal AI company Harvey on model-harness co-optimization for end-to-end M&A diligence: applying RL post-training to a Qwen 122B-A10B root agent inside an RLM harness lifted the pass rate from 29% to 63%.
- The RLM (recursive language model) harness lets the agent decompose long-horizon legal workflows into tool calls and subtasks
- Key takeaway: swapping models alone isn't enough—co-optimizing model and harness for the task drives large end-to-end gains
- The team is now scaling RL training with GLM-5.3, a frontier open-weight model, and observing similar uplift
- This extends Harvey's Tenet research preview toward its legal agent products
Related event: Baseten and Harvey boost legal agent due diligence with RL-trained RLM(2 posts)→
More from coding & agent
- 'Spawning 100,000 sub agents': the meme about agents over-refactoring on demand — NERDDISCO · 2026-09-09
- Early Astra agent test flops: free-rein CUDA kernel optimization fails on architecture — gandamu_ml · 2026-09-09
- Anthropic's Claude Tag acts as on-call first responder for CI/CD failures — ClaudeDevs · 2026-09-09
- Agents now ship with a soul.md file, a nod to Peter Steinberger's influence — altryne · 2026-09-09
- How do you coordinate 10,000 AI agents on one problem? Hierarchical orchestration ideas emerge — pwlot · 2026-09-09
- Mitchell Hashimoto demos Superlogical remote persistent sessions, a full SSH replacement — iannuttall · 2026-09-09