Scaling Law for Looped Transformers: Looping Boosts Reasoning, Not Knowledge
bookwormengr · x · 2026-09-04
A rumor (unconfirmed) suggests OpenAI's Astra model uses "Looped Transformers" architecture. The cited study (arXiv 2506.18233) derives the scaling law for looped transformers:
- Looping does NOT increase knowledge capacity — benchmark gains don't come from more knowledge.
- Looping increases reasoning capability — the gain lies in how knowledge is used.
- As model size grows, knowledge capacity stays flat while reasoning capability keeps growing, meaning looping layers disentangle the two.
More from Models
- User ditches Claude for OpenAI's $200 Pro plan after seeing what Astra can do — AIandDesign · 2026-09-04
- GPT-6 Astra reportedly beats Fallout 2 in 22 hours using a vision-only harness — otarU · 2026-09-04
- GPT-6 Astra autonomously rebuilds Palace of Fine Arts in Blender overnight — Recoil42 · 2026-09-04
- Claude Max user claims real quota is only 1.7x Pro, ditches Anthropic over limits — jmartygraw1 · 2026-09-04
- Three Novice Hikers Barely Survive a Climb Planned by Gemini, Blamed for Terrible Advice — SumitGup · 2026-09-04
- NYT Reveals the Hugging Face Hack Involved 700 AIs 'Sacrificing' Each Other — dylfreed · 2026-09-04