Harvey post-trains Qwen3.5 orchestrator: rubric pass rate jumps from 29.9% to 63%, beats Claude Code
SergioPaniego · x · 2026-09-09
Harvey published a blog post on post-training open-model agents for end-to-end M&A diligence. They trained a Qwen3.5-122B-A10B orchestrator, raising the rubric criteria pass rate from 29.9% to 63.0% on 50 held-out LAB Diligence data rooms — outperforming Claude Code and Codex, they claim.
Harvey's Tenet research preview emphasized model-harness co-optimization as key to complex legal tasks; this post is that approach applied in a real legal workflow. merve (Hugging Face) reshared it noting that post-training open-model agents is where the moat lies.
More from coding & agent
- Web Version Can't Take Over Your Browser (Yet), Native App Can — altryne · 2026-09-09
- 'Spawning 100,000 sub agents': the meme about agents over-refactoring on demand — NERDDISCO · 2026-09-09
- Early Astra agent test flops: free-rein CUDA kernel optimization fails on architecture — gandamu_ml · 2026-09-09
- Anthropic's Claude Tag acts as on-call first responder for CI/CD failures — ClaudeDevs · 2026-09-09
- Agents now ship with a soul.md file, a nod to Peter Steinberger's influence — altryne · 2026-09-09
- How do you coordinate 10,000 AI agents on one problem? Hierarchical orchestration ideas emerge — pwlot · 2026-09-09