New Paper Proposes Recursive Transformers for Model Compression via Layer Sharing

max_paperclips · x · 2026-09-02

A new paper explores compressing large models via layer parameter sharing, termed "Relaxed Recursive Transformers." Built on pretrained Transformers, this method loops a single block of unique layers and adds depth-wise LoRA adapters for flexibility. Experiments show it matches or outperforms similar-sized models like TinyLlama and recovers most performance of the full-size model. Discussions also touch on "depth-wise batching" for better utilization, noting the main advantages are storage and KV cache optimization rather than raw speed.

Original post →

More from Infra

Infra channel →