Harnesses widen what LMs can do, but may not improve compositional generalization
a1zhang · x · 2026-07-23
The post argues that harnesses help expand the set of tasks LMs can solve, but not necessarily through true compositional generalization.
It says code settings rely on tools such as greppers, code tools, and skills mainly because the raw LM interface is too limited. Even when a harness makes tasks easier inside a domain, the model still depends on its own generalization to handle new cases. The author also stresses that comparing a harnessed setup to a base Transformer is useful for measuring the lift from training around a harness, and for clarifying what the LM inside the harness actually sees.
More from coding & agent
- Cursor Launches Intelligent Router: Dynamic Model Selection Cuts Costs by 60% — stuffyokodraws · 2026-07-23
- A Codex Micro user swaps push-to-talk for Wispr Flow and gets a faster workflow — Dimillian · 2026-07-23
- DeepMind’s 180-agent study says teams win on split tasks, but lose on sequential work — brucemacv · 2026-07-23
- FiveClaw adds a managed MCP codespace for FiveM AI development — nytro_Haze · 2026-07-23
- Custom Anthropic chat wrapper uses citations, voice commands, and costs about $30 a month — BLUECOW009 · 2026-07-23
- Migma: A Cross-Platform Email Rendering Engine Built for AI Agents — bigaiguy · 2026-07-23