Are URM and Universal Transformers the forgotten architecture beating standard LLMs?
moschles · reddit · 2026-10-09
- Redditor moschles asks whether Universal Transformers (UT) and Universal Reasoning Models (URM) have been quietly absorbed into frontier models, or left as forgotten papers.
- UT replaces stacked layers with recurrent computation over depth, repeatedly applying one transition block (H^(t+1) = LayerNorm(H^t + MHA(H^t)) plus a position-wise Transition), with 2D sinusoidal embeddings encoding both position and refinement depth.
- The cited paper (arXiv 2512.14693) claims UT-based small models, trained from scratch without internet-scale pretraining, consistently outperform most standard Transformer LLMs by a significant margin. URM adds fixed loops, ACT loops, a ConvSwiGLU module, and a novel Truncated Backpropagation Through Loops scheme.
- The post links the paper, a simplified blog, and a YouTube talk, inviting discussion on whether big labs already use these enhancements.
More from Models
- ValsAI: Inception's Mercury Decide matches frontier accuracy at the lowest cost measured — pratyusha_PS · 2026-10-09
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09