Telescopic LM trains one model valid at every depth, cutting quality-budget area 43%

iScienceLuvr · x · 2026-09-29

A new arXiv paper introduces Telescopic Language Models (TLM): a nested-capacity Transformer trained with stochastic prefix supervision plus a full anchor. Each step runs two forward-backward passes—one on a randomly truncated prefix of the capacity axis against the full next-token target, one at full capacity—with no architectural change and zero inference overhead, making the model valid at any depth.

The authors contrast fixed-exit suites like MLMS, where supervising only a few exits leaves other depths at chance level (perplexity 10^2–10^5 in baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data), a single TLM run is a valid LM at all twenty layer prefixes in perplexity and perplexity-sensitive tasks, reducing the quality-budget curve area by 43–44% versus fixed-exit suites while matching full capacity, at 12% lower GPU cost per run. Prefix sampling density is a dial: concentrating it recovers fixed-exit training.

Original post →

More from Research

Research channel →