New LLM papers: recursive language models generalize out of domain; XMerge depth compression

burny_tech · x · 2026-09-23

Two LLM research shares: Recursive Language Models Generalize Out of Domain (Toyota Technological Institute at Chicago; Yang, Li, McAllester, Srebro) on out-of-domain generalization of recursive LMs, and XMerge (arXiv:2609.02083), a post-training depth-compression method combining cross-axis block selection with local boundary reconstruction. Across seven Llama/Qwen backbones (0.5B–8B), XMerge ranks first on six of seven on CORE and MMLU at k=4 layer removal, avoids large perplexity spikes, and is the only operator that never collapses across all 14 (model, regime) cells — with no task labels, fine-tuning, or architectural changes.

Original post →

More from Infra

Infra channel →