Are Transformers really deeper-efficient than RNNs? Paper questions layer-count advantage
kfountou · x · 2026-09-21
With polynomial-size MLPs, it's unclear whether Transformers hold any advantage over RNNs in terms of layer count, the author argues, linking a paper on architecture depth-efficiency comparisons.
Related event: Flaw Claimed in Transformer Logarithmic Depth Proof(3 posts)→
More from Research
- GaME (CVPR 2026): Gaussian mapping that forgets stale geometry as robot scenes change — lucacarlone1 · 2026-09-21
- YOCO back in spotlight: blog breaks down cross-layer KV sharing in DeepSeek-V4.1-Flash and Gemma 4 — donglixp · 2026-09-21
- Dead Human Brain Tissue Controls Robot: Hybrots Are 20 Years Old, Argue Critics — ryunuck · 2026-09-21
- New paper: simple difference-of-means vectors detect reward hacking from LLM internals before it happens — burny_tech · 2026-09-21
- What AI means for mathematicians: seven predictions extrapolated from software — cgarciae88 · 2026-09-21
- Qwen releases RecreationBench: 250 tasks testing agents that rebuild real apps — burny_tech · 2026-09-21