Transformers with Memory and Recursion Are No Longer Pure DNNs

lateinteraction · x · 2026-08-06

Responding to recent community mockery of LeCun and Marcus for criticizing "pure LLMs," the author offers a rigorous technical perspective: modern Transformers equipped with scratchpads, offloaded context, PTC, and recursion are unequivocally no longer pure Deep Neural Networks (DNNs).

This architectural evolution is precisely why they exceed the capabilities of early models like plain GPT-4 with RLHF. The author argues for clearer distinctions in our hypothesis classes, separating system-level evolution from static network architectures.

Related event: LLM Architecture Debate: Equipped Transformers No Longer Pure DNNs(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →