Agentic infra can beat LLMs alone on compositional generalization

prasanna_says · x · 2026-07-21

The post argues that the output of **agentic infrastructure** is greater than the sum of an LLM plus its harness. It connects this idea to **compositional generalization**: if a model can solve one task, it does not necessarily generalize to a structurally similar unseen task. A harness can help by making two problems look the same through context offloading and programmatic sub-calling, which may let models trained on easy tasks generalize better to harder ones. The author suggests this pattern matters not only for closed-loop domains like coding and math, but also for other decomposable tasks where humans may currently provide the missing loop via taste and selection of outputs.

Related event: AI Agent Architecture Reflections: Thinner Harnesses and Multi-Agent Tradeoffs(13 posts)→

Original post →

More from coding & agent

coding & agent channel →