Why Hasn't Anyone Tested Replacing Most Attention Layers with ResNet MLPs?

kalomaze · x · 2026-07-30

Developer @kalomaze raised a counter-intuitive question about model architecture: while the industry explores linear attention or Mamba-like RNN variants, almost no one has tested a trivial baseline—using basic ResNet MLPs for over 80% of the layers while keeping softmax attention only for the rest.

Related event: Researcher Questions Linear Attention with Basic MLP Baseline(4 posts)→

Original post →

More from Research

Research channel →