Tsinghua and ByteDance Seed unveil SMELT, the first fair comparison of Looped Transformers
jiqizhixin · x · 2026-09-18
Looped Transformers repeat the same layer for deeper computation without new parameters, showing promise on reasoning and math — but is the gain from the architecture or just more FLOPs? Tsinghua University and ByteDance Seed present SMELT, the first fair comparison across multiple model scales:
- A fair comparison must control per-token FLOPs (training/inference cost), total parameters (knowledge capacity), and KV cache (servable context length) simultaneously
- SMELT uses MoE to control parameter count while matching FLOPs and KV cache, comparing Looped vs standard Transformers at truly equal budget
- First author: Sha… (truncated in source)
More from Research
- Who Gets Credit When AI Proves Collatz? The New Math Attribution Dilemma — jd_pressman · 2026-09-18
- SceneAgent: agentic pipeline turns 3D captures into physics-ready scenes for robot training — hankyang94 · 2026-09-18
- LAX launches to bridge natural math language and Lean, but looks a lot like existing tool span — lpachter · 2026-09-18
- Fari Research paper: misaligned AI may just persuade its human overseers — DG_Rand · 2026-09-18
- Biomedical world models: a framework for virtual drug trials, intervention design and planning — marinkazitnik · 2026-09-18
- LeanReact 0.1: expressing composable, provably correct React components in Lean — hargup13 · 2026-09-18