Telescopic LM trains one model valid at every depth, cutting quality-budget area 43%
iScienceLuvr · x · 2026-09-29
A new arXiv paper introduces Telescopic Language Models (TLM): a nested-capacity Transformer trained with stochastic prefix supervision plus a full anchor. Each step runs two forward-backward passes—one on a randomly truncated prefix of the capacity axis against the full next-token target, one at full capacity—with no architectural change and zero inference overhead, making the model valid at any depth.
The authors contrast fixed-exit suites like MLMS, where supervising only a few exits leaves other depths at chance level (perplexity 10^2–10^5 in baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data), a single TLM run is a valid LM at all twenty layer prefixes in perplexity and perplexity-sensitive tasks, reducing the quality-budget curve area by 43–44% versus fixed-exit suites while matching full capacity, at 12% lower GPU cost per run. Prefix sampling density is a dial: concentrating it recovers fixed-exit training.
More from Research
- Economist John Horton: keep AI-generated papers out of human venues, publish on GitHub — soumitrashukla9 · 2026-09-29
- Harvard's Clinical Informatics Lecture Series hosts "AI and Mental Health" with Dr. John Torous — zakkohane · 2026-09-29
- HCOMP 2026 study: imagery scales skew annotator and reviewer performance — windx0303 · 2026-09-29
- EleutherAI open-sources Bergson, a scalable data attribution library for LLMs, with an EMNLP 2026 oral paper — BlancheMinerva · 2026-09-29
- Bergson data attribution library released with first public MAGIC implementation — BlancheMinerva · 2026-09-29
- EleutherAI tests whether EK-FAC data attribution can block subliminal learning — results are 'solidly mid' — BlancheMinerva · 2026-09-29