Researcher Questions Linear Attention, Suggests Sparse MLP Baseline

Independent researcher kalomaze questioned the necessity of new architectures like linear attention, noting that the stability of gradient descent is more crucial than the function class itself. He suggested that using basic sparse MLPs or ResNets to replace 90% of attention layers could be a surprisingly strong, yet underexplored baseline.

2026-07-30 ~ 2026-07-30 · 4 related posts