Harness tuning lifts Qwen3.6-27B legal agent pass rate from 67% to 85%, study finds

sarahookr · x · 2026-09-09

Research shared by sarahookr argues agent performance is wrongly attributed mostly to the underlying model: the harness can only recover what a model can already do, not create capability it lacks.

Key numbers: on the Harvey Legal Agent Benchmark, harness optimization alone moved Qwen3.6-27B's pass rate from 67% to 85%, with post-training pushing it only to 88%.

The author's framing: optimize the entire stack, like a chef wanting both the best ingredients (model) and the best oven (harness), rather than fixating on one.

Original post →

More from coding & agent

coding & agent channel →