Questioning the Baseline: What If 90% of LLM Layers Were Just Basic MLPs?
kalomaze · x · 2026-07-30
AI researcher kalomaze raised an interesting architectural question: while the industry explores linear attention or Mamba-esque RNNs, almost no one seems to have tested a trivial baseline. What if we used basic ResNet MLPs for over 90% of the layers and simply let the residual stream mix the information?
Related event: Researcher Proposes Minimal MLP Baseline to Replace Most Attention Layers(5 posts)→
More from Research
- When Does Synthetic Data Work? Research Reveals Optimal Ratios and 'Zeta Law' — PTenigma · 2026-07-30
- Applying Jacobian Methods for LLM Contrastive Steering Outperforms Controls — voooooogel · 2026-07-30
- Quadratic Models Surprisingly Accurately Describe LLM Pretraining, Paper Finds — jasondeanlee · 2026-07-30
- Lean as the Ultimate Echo of Principia Mathematica: A Philosophical Divide — doodlestein · 2026-07-30
- Meta & CMU Paper: Agentic Context Management Boosts Long-Horizon Task Performance by 27% — rohanpaul_ai · 2026-07-30
- Scaling Semiconductor Quantum Computers: Qubits Need to Match Classical Transistors — whurley · 2026-07-30