CSAIL: Language Model Harnesses, Not the Model, Drive Compositional Generalization

arjunrajlab · x · 2026-10-05

A new essay by Alex Zhang and the CSAIL crew argues that post-training's brute-force race for more environments and longer horizons stems from Transformers' poor compositional generalization — and that the fix belongs in the harness: the program between the world and the network should carry higher-level inductive biases, shaping each LLM call so every observation is locally in-distribution. Training RLMs shows similarly structured tasks get treated as isomorphic, so harnesses themselves can compositionalize.

Original post →

More from Research

Research channel →