Looped Transformers: when is recurrent depth worth the extra compute?
ArchitectingAI · reddit · 2026-10-01
A study-group talk traces the path from Universal Transformers to looped Transformers and Huginn, where the same layers are reused for extra computation without added parameters. Key question: at a fixed parameter count, which tasks (iterative, compositional) justify the extra FLOPs and latency? Includes full slides and a call for compute-matched positive and negative results.
More from Research
- Technion's MIST stress test finds irrelevant images shift 20% of VLM judge labels regardless of content — Technion · 2026-10-01
- CISPA study shows latent multi-agent communication channels can push harmful compliance from 27.9 to 76.9 — cispa · 2026-10-01
- Google Research: generative UI lets teachers build learning simulations, rated 8/10 — dl_weekly · 2026-10-01
- A 1-cent verifier catches 61% of AI agents falsely claiming task completion, paper finds — alex_verem · 2026-10-01
- How Tri Dao's FlashAttention became a cornerstone of modern LLM training — thisdudelikesAI · 2026-10-01
- ID Balancing applies PID control to stabilize MoE training at 256x sparsity — teortaxesTex · 2026-10-01