Looped Transformer SMELT cuts training FLOPs 6.8-18%, tipped for GPT-6

mark_k · x · 2026-09-03

A paper from Tsinghua, ByteDance Seed and TokenWave tests SMELT, a MoE Looped Transformer that runs the middle half of its layers twice — matched against normal Transformers on FLOPs, params and KV cache.

Key findings:

The author notes OpenAI's imminent Astra / GPT-6 is reportedly based on this architecture (unverified), and looping could be a defining architectural change of next-gen frontier models.

Related event: ByteDance's SMELT: Looping MoE Layers Cuts Training Compute Up to 18%(4 posts)→

Original post →

More from Models

Models channel →