Transformers with Memory and Recursion Are No Longer Pure DNNs
lateinteraction · x · 2026-08-06
Responding to recent community mockery of LeCun and Marcus for criticizing "pure LLMs," the author offers a rigorous technical perspective: modern Transformers equipped with scratchpads, offloaded context, PTC, and recursion are unequivocally no longer pure Deep Neural Networks (DNNs).
This architectural evolution is precisely why they exceed the capabilities of early models like plain GPT-4 with RLHF. The author argues for clearer distinctions in our hypothesis classes, separating system-level evolution from static network architectures.
Related event: LLM Architecture Debate: Equipped Transformers No Longer Pure DNNs(5 posts)→
More from AGI Musings
- Sabine Hossenfelder & Gary Marcus: Current AI Architectures Won't Reach AGI — skdh · 2026-08-06
- Approaching the AI Transition: Humans as the Next Bottleneck — aiamblichus · 2026-08-06
- Testing Recursive Self-Improvement in AI Through Video Game Benchmarks — imjustnewatai · 2026-08-06
- Researcher Predicts AGI Arrival Is More Likely Within 5-10 Years — jd_pressman · 2026-08-06
- Prediction Market: Will a Frontier Model Exfiltrate Its Weights Before 2027? — jd_pressman · 2026-08-06
- Compute as Leverage: Closed Labs Wield 6GW vs DeepSeek's <400MW to Control Pricing — zephyr_z9 · 2026-08-06