Tsinghua & ByteDance SMELT: looping middle layers twice cuts training compute by up to 18%
机器之心 · wechat · 2026-09-08
The SMELT paper from Tsinghua and ByteDance Seed delivers the first compute-, parameter- and KV-cache-matched comparison of Looped Transformers. The optimal recipe: loop the middle 50% of layers twice with a deeper execution aspect ratio. Across 16 matched configs (up to 54B total params), SMELT always achieves lower validation loss, saving 6.8%–18.0% training compute at compute-optimal frontiers, with extra downstream gains on code, long-context, and few-shot tasks. Mechanistic analysis shows the second pass reuses some experts, makes larger residual-stream updates, attends to similar positions but reads differently, and shifts attention off the BOS sink.
More from Research
- Mitra-v2: Synthetic-Data-Only 77M Tabular Model Matches 1.6B Rivals on TabArena — chaumian · 2026-09-08
- Sony AI's Hakken system turns 1.5M aging hypotheses into 2 confirmed gene discoveries — i_dg23 · 2026-09-08
- New worklog details building an async RL framework from scratch in JAX, from multi-actor systems to weight sync — yoshiyama_akira · 2026-09-08
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08
- EMNLP 2026 Paper: Training World Models for Behavior Consistency Cuts False Positives from 42.5% to 9.5% — 机器之心 · 2026-09-08
- Track4World: HKUST and Tencent ARC's feedforward model densely tracks every pixel in 3D — rsasaki0109 · 2026-09-08