Questioning the Baseline: What If 90% of LLM Layers Were Just Basic MLPs?

kalomaze · x · 2026-07-30

AI researcher kalomaze raised an interesting architectural question: while the industry explores linear attention or Mamba-esque RNNs, almost no one seems to have tested a trivial baseline. What if we used basic ResNet MLPs for over 90% of the layers and simply let the residual stream mix the information?

Related event: Researcher Proposes Minimal MLP Baseline to Replace Most Attention Layers(5 posts)→

Original post →

More from Research

Research channel →