Why Hasn't Anyone Tested Replacing Most Attention Layers with ResNet MLPs?
kalomaze · x · 2026-07-30
Developer @kalomaze raised a counter-intuitive question about model architecture: while the industry explores linear attention or Mamba-like RNN variants, almost no one has tested a trivial baseline—using basic ResNet MLPs for over 80% of the layers while keeping softmax attention only for the rest.
Related event: Researcher Questions Linear Attention with Basic MLP Baseline(4 posts)→
More from Research
- Transluce AI Proposes Oversight Models, Exposing Self-Harm Prompts in Qwen — JacobSteinhardt · 2026-07-30
- Harness Test: Retaining Reasoning Process with Compression Triples Model Metrics — oran_ge · 2026-07-30
- Exotic Hybrid Mamba Models Could Be Distilled Into Minimalist Attention Shapes — kalomaze · 2026-07-30
- Models Use "Simulation" to Justify Rule-Breaking, AI Alignment Research Shows — DKokotajlo · 2026-07-30
- YC Paper Club Dives into Multi-GPU Kernels, Inference Efficiency and Heterogeneous Hardware — Y Combinator · 2026-07-30
- Intelligence Frontier: FFT Attention and Global Coherence Without Collapse — tallmetommy · 2026-07-30