Harnesses may be widening LM capability without creating true generalization
a1zhang · x · 2026-07-23
The post argues that the harness around an LM can reduce the burden of generalization: instead of relying only on the model’s internal mechanisms to “grok everything,” the setup can make the model see at test time the same kinds of things it saw during evaluation.
A reply pushes back on the stronger claim that harnesses help generalization in the compositional sense. The counterpoint is that harnesses likely expand the class of tasks LMs can solve, especially in code-like settings, but that success may come from scaffolding rather than true generalization.
More from Research
- Why the FAccT Conference Became an Unlikely Powerhouse in AI Policy — rajiinio · 2026-07-23
- Raji says FAccT papers are heavily cited in NIST, FTC and DOJ AI policy docs — rajiinio · 2026-07-23
- Nature piece says LLMs can forecast social-science experiment outcomes — RobbWiller · 2026-07-23
- Nature paper finds LLMs can predict social-science experiment results — RobbWiller · 2026-07-23
- Researchers open a demo for forecasting social-science treatment effects — RobbWiller · 2026-07-23
- LLM-only pilots cost under $1 and rival ~230-person human pilots — RobbWiller · 2026-07-23