Researchers argue Transformer encoder/decoder naming is just attention masks

TimDarcet · x · 2026-09-10

In a discussion sparked by "encoder-decoder is so back," jramapuram argued the encoder/decoder naming in Transformers is misleading: everything reduces to self- or cross-attention with different masks, unlike VAEs where the naming reflects an actual bottleneck and learned posteriors. A concise technical take amid renewed interest in encoder-decoder architectures.

Related event: Researchers call encoder/decoder Transformer naming a misnomer: it's all attention masks(2 posts)→

Original post →

More from Research

Research channel →