Tsinghua & ByteDance SMELT: looping middle layers twice cuts training compute by up to 18%

机器之心 · wechat · 2026-09-08

The SMELT paper from Tsinghua and ByteDance Seed delivers the first compute-, parameter- and KV-cache-matched comparison of Looped Transformers. The optimal recipe: loop the middle 50% of layers twice with a deeper execution aspect ratio. Across 16 matched configs (up to 54B total params), SMELT always achieves lower validation loss, saving 6.8%–18.0% training compute at compute-optimal frontiers, with extra downstream gains on code, long-context, and few-shot tasks. Mechanistic analysis shows the second pass reuses some experts, makes larger residual-stream updates, attends to similar positions but reads differently, and shifts attention off the BOS sink.

Original post →

More from Research

Research channel →