Researcher Questions Linear Attention, Suggests Sparse MLP Baseline
Independent researcher kalomaze questioned the necessity of new architectures like linear attention, noting that the stability of gradient descent is more crucial than the function class itself. He suggested that using basic sparse MLPs or ResNets to replace 90% of attention layers could be a surprisingly strong, yet underexplored baseline.
2026-07-30 ~ 2026-07-30 · 4 related posts
- Questioning the Baseline: What If 90% of LLM Layers Were Just Basic MLPs? — kalomaze · 2026-07-30
- Why Hasn't Anyone Tested Replacing Most Attention Layers with ResNet MLPs? — kalomaze · 2026-07-30
- Researcher Questions Linear Attention: Sparse MLPs May Be a Stronger Baseline — kalomaze · 2026-07-30
- Architectural Innovation: Gradient Descent Stability Trumps Function Class — kalomaze · 2026-07-30