Fully Nested Transformers: A New Architecture

adityakusupati · x · 2026-07-10

A post highlights a new architecture called Fully Nested Transformers. By utilizing block triangular weights, it nests an entire family of sub-models within a single model, supporting capabilities like token-adaptive routing and self-distillation inference. The original post emphasizes that this makes it possible to "train a family of sub-models" and "use them simultaneously."

Original post →

More from Research

Research channel →