Harness tuning lifts Qwen3.6-27B legal agent pass rate from 67% to 85%, study finds
sarahookr · x · 2026-09-09
Research shared by sarahookr argues agent performance is wrongly attributed mostly to the underlying model: the harness can only recover what a model can already do, not create capability it lacks.
Key numbers: on the Harvey Legal Agent Benchmark, harness optimization alone moved Qwen3.6-27B's pass rate from 67% to 85%, with post-training pushing it only to 88%.
The author's framing: optimize the entire stack, like a chef wanting both the best ingredients (model) and the best oven (harness), rather than fixating on one.
More from coding & agent
- Tracking agent activity is a headache: how do you get real observability in production? — ComparisonNew9425 · 2026-09-09
- aiden bai is net bearish on runtime MCP: Codex Computer Use / Playwright suffice — aidenybai · 2026-09-09
- Simple Post ChatGPT plugin approved: write and schedule social posts inside your chat — haltakov · 2026-09-09
- LangChain Shows 3-Minute Workflow to Turn Flagged Traces Into Eval Datasets via LangSmith CLI — LangChain · 2026-09-09
- Anybrowse launches MCP-native scraping API with 90% success rate on Cloudflare-protected sites — modelcontextprotocol · 2026-09-09
- URL Safety Validator MCP rates links SAFE/SUSPICIOUS/DANGEROUS with trust scores — modelcontextprotocol · 2026-09-09