New LLM papers: recursive language models generalize out of domain; XMerge depth compression
burny_tech · x · 2026-09-23
Two LLM research shares: Recursive Language Models Generalize Out of Domain (Toyota Technological Institute at Chicago; Yang, Li, McAllester, Srebro) on out-of-domain generalization of recursive LMs, and XMerge (arXiv:2609.02083), a post-training depth-compression method combining cross-axis block selection with local boundary reconstruction. Across seven Llama/Qwen backbones (0.5B–8B), XMerge ranks first on six of seven on CORE and MMLU at k=4 layer removal, avoids large perplexity spikes, and is the only operator that never collapses across all 14 (model, regime) cells — with no task labels, fine-tuning, or architectural changes.
More from Infra
- Cloudflare Bets on Microtransactions to Save the Web, but Agents Want Answers, Not Fragments — Ronangmi · 2026-09-23
- Google Colab joins Google AI plans with faster GPUs and background execution — DynamicWebPaige · 2026-09-23
- NEAR AI confidential inference goes live on Bittensor via SayGm's TDX-enclave routing — markjeffrey · 2026-09-23
- 6 serving-side techniques that make LLM inference faster - from prefix caching to PD disaggregation — techNmak · 2026-09-23
- SanDisk shares surge 7.56% as Wall Street turns bullish on AI memory demand — Polymarket · 2026-09-23
- SGLang v0.5.20 ships: Intel XPU support, up to 52% faster decode — hsu_byron · 2026-09-23