CSAIL: Language Model Harnesses, Not the Model, Drive Compositional Generalization
arjunrajlab · x · 2026-10-05
A new essay by Alex Zhang and the CSAIL crew argues that post-training's brute-force race for more environments and longer horizons stems from Transformers' poor compositional generalization — and that the fix belongs in the harness: the program between the world and the network should carry higher-level inductive biases, shaping each LLM call so every observation is locally in-distribution. Training RLMs shows similarly structured tasks get treated as isomorphic, so harnesses themselves can compositionalize.
More from Research
- Jon Barron walks through backpropagation by hand on a tiny two-layer network — techNmak · 2026-10-06
- New method dissects only task-relevant weights, making interpretability cheap enough for daily debugging — CatAstro_Piyush · 2026-10-06
- Debating computational irreducibility: if you've computed the Mandelbrot set, is the program just compression? — ctjlewis · 2026-10-06
- Social media use explains just 0.4% of teen well-being variation, researcher argues studies fail policy — asusarla · 2026-10-06
- Stanford Open-Sources DITTO-X: Force-Feedback Teleop With Reverse Human Intervention — CyberRobooo · 2026-10-06
- Cisco Benchmarks Decision Models: Jev Nears 31B LLM Judge on Zero-Shot Safety Classification — aminkarbasi · 2026-10-06