Looped Transformer hype: Nanbeige4.2-3B beats 12B models on agent benchmarks

alexcovo_eth · x · 2026-09-08

After The Information reported that OpenAI's GPT-6 Astra uses a "recurrent depth or looped transformer", rasbt published a debunking thread explaining the idea is simply reusing the layer stack within transformer blocks to add capacity without parameters — pointing to Nanbeige, which disclosed the same architecture two months ago: Nanbeige4.2-3B, pretrained from scratch on 28T tokens.

The Nanbeige team followed up claiming the 3B model consistently outperforms larger models (Qwen3.5-9B, Gemma4-12B) on agent benchmarks including SWE-Bench Pro, Terminal-Bench and GDPval. Nanbeige4.5 is in training with LoopSplit, mHC + depth attention, and concatenated n-gram embeddings.

Original post →

More from Models

Models channel →