Harnesses widen what LMs can do, but may not improve compositional generalization

a1zhang · x · 2026-07-23

The post argues that harnesses help expand the set of tasks LMs can solve, but not necessarily through true compositional generalization.

It says code settings rely on tools such as greppers, code tools, and skills mainly because the raw LM interface is too limited. Even when a harness makes tasks easier inside a domain, the model still depends on its own generalization to handle new cases. The author also stresses that comparing a harnessed setup to a base Transformer is useful for measuring the lift from training around a harness, and for clarifying what the LM inside the harness actually sees.

Related event: Researchers Debate: Does LLM Generalization Come from the Model or the Harness?(8 posts)→

Original post →

More from coding & agent

coding & agent channel →