RLM harness lifts M&A diligence pass rate from 23.3% to 62.4% across seven models
a1zhang · x · 2026-09-09
- nikogrupen, with Baseten and Baselabs, built an RLM harness for M&A diligence: a root agent delegates document review to sub-agents and aggregates their findings.
- Harness design: across seven models, moving from a standard tool loop to the RLM harness raised mean criteria pass rate from 23.3% to 62.4%.
- Post-training: in a separate experiment, RL on Qwen3.5 inside the RLM harness more than doubled rubric pass rate, starting from 29.9%.
- The results show model-harness co-optimization delivering concrete gains on end-to-end legal workloads.
Related event: Baseten and Harvey boost legal agent due diligence with RL-trained RLM(2 posts)→
More from coding & agent
- LangChain details subagent forking as context engineering trick in deepagents — LangChain · 2026-09-09
- Box launches Mount to sync Box folders into agent sandboxes with built-in governance — badphilosopher · 2026-09-09
- Anthropic interviews WisprFlow, Actively and Pendo on building with Claude Managed Agents — ClaudeDevs · 2026-09-09
- Using ChatGPT Sites as a progress log for long-running agent projects — jdjohnson · 2026-09-09
- Muse review: agentic AI for normies with UI better than Anthropic or OpenAI apps — neil_chilson · 2026-09-09
- mumo MCP server routes questions across Claude, GPT, Gemini, Grok for cross-model debate — modelcontextprotocol · 2026-09-09