Looped Transformer hype: Nanbeige4.2-3B beats 12B models on agent benchmarks
alexcovo_eth · x · 2026-09-08
After The Information reported that OpenAI's GPT-6 Astra uses a "recurrent depth or looped transformer", rasbt published a debunking thread explaining the idea is simply reusing the layer stack within transformer blocks to add capacity without parameters — pointing to Nanbeige, which disclosed the same architecture two months ago: Nanbeige4.2-3B, pretrained from scratch on 28T tokens.
The Nanbeige team followed up claiming the 3B model consistently outperforms larger models (Qwen3.5-9B, Gemma4-12B) on agent benchmarks including SWE-Bench Pro, Terminal-Bench and GDPval. Nanbeige4.5 is in training with LoopSplit, mHC + depth attention, and concatenated n-gram embeddings.
More from Models
- Astra is 'actually quite good' at using a browser, though still not fast — jobergum · 2026-09-08
- Codex tip: Astra now reads your remaining usage % so you can budget in plain English — Dimillian · 2026-09-08
- Matt Shumer calls GPT-6 Pro "a monster" in unverified hype tweet — msg · 2026-09-08
- "The old models just sucked": stronger models redefine how intensely you use AI — teortaxesTex · 2026-09-08
- Astra announces global usage reset for all paid subscriptions today — infoxiao · 2026-09-08
- New Codex 7-day limit visualization shows red when you're running ahead of your quota — lucasmeijer · 2026-09-08