Looped Transformers work best from scratch, with two passes emerging as the sweet spot

jm_alexia · x · 2026-07-27

Key takeaways from a loop-architecture tech report

The quoted summary highlights several ablation findings about looped Transformers:

The context mentions Nanbeige4.2-3B, a looped Transformer released as a capable 3B agent, and says Nanbeige4.5 is already training with LoopSplit, mHC+depth attention, and concatenated n-gram embeddings.

Original post →

More from Models

Models channel →