Agentic infra can beat LLMs alone on compositional generalization
prasanna_says · x · 2026-07-21
The post argues that the output of **agentic infrastructure** is greater than the sum of an LLM plus its harness. It connects this idea to **compositional generalization**: if a model can solve one task, it does not necessarily generalize to a structurally similar unseen task. A harness can help by making two problems look the same through context offloading and programmatic sub-calling, which may let models trained on easy tasks generalize better to harder ones. The author suggests this pattern matters not only for closed-loop domains like coding and math, but also for other decomposable tasks where humans may currently provide the missing loop via taste and selection of outputs.
More from coding & agent
- OxDeAI opensource protocol moves AI-agent policy checks before execution — docybo · 2026-07-21
- Kimi staff member builds a VR companion with Kimi Code K3 demo — dejavucoder · 2026-07-21
- Kimi K3 took 75 minutes and still failed a simple diagram task, user says — MinusKarma01 · 2026-07-21
- Coding agents need better rules for when to read search summaries or full pages — RhubarbLarge2747 · 2026-07-21
- Chart compares open tickets across Gemini, Codex, Claude, Grok and Opus 3 — nptacek · 2026-07-21
- Notch says he may try vibe coding after struggling to hire good programmers — max_paperclips · 2026-07-21