Non-transformer model with 150M params redraws ARC-AGI cost frontier, validated by Transformer co-author
Ok_Can_1968 · reddit · 2026-08-14
A new study shows a non-transformer architecture achieving state-of-the-art on ARC-AGI-1 with only 150M parameters and drastically lower compute, while retaining recurrent latent reasoning at 600B scale. Evaluated by Transformer co-author Łukasz Kaiser and others, the approach uses recurrent memory and latent reasoning instead of brute-force scaling, potentially reshaping AI economics.
More from AGI Musings
- Elon Musk says wealth distribution "won't be relevant in the future" — critics push back — StewartalsopIII · 2026-08-14
- Ken Griffin: AI Toolkit Profoundly More Powerful, Man-Years of Work Done in Days — damianplayer · 2026-08-14
- AI models need to stay current; future learning will happen outside parameters — coallaoh · 2026-08-14
- Post-training works but is gated; updating parameters risks existing knowledge — coallaoh · 2026-08-14
- Model adaptation happens outside parameters: engineering, memory, RAG, tools — coallaoh · 2026-08-14
- Bet: most learning keeping models current will happen outside parameters — coallaoh · 2026-08-14