Researchers argue Transformer encoder/decoder naming is just attention masks
TimDarcet · x · 2026-09-10
In a discussion sparked by "encoder-decoder is so back," jramapuram argued the encoder/decoder naming in Transformers is misleading: everything reduces to self- or cross-attention with different masks, unlike VAEs where the naming reflects an actual bottleneck and learned posteriors. A concise technical take amid renewed interest in encoder-decoder architectures.
More from Research
- Bryan Johnson responds to Michael Levin's peer-reviewed Platonic Space paper: bodies as collective intelligence — AllThingsApx · 2026-09-10
- OpenBMB open-sources training data: 400B tokens of code and 500K agent samples — zibuyu9 · 2026-09-10
- RobustSGPO lifts agent harness completion from 60% to 80% on held-out tasks — dair_ai · 2026-09-10
- Jack Clark proposes pre-registering AI economy forecasts to score predictors in a year — jackclarkSF · 2026-09-10
- TokenPrint: open-source interactive explorer for tokens, attention, hidden states and KV cache seeks contributors — Rich-Fruit-326 · 2026-09-10
- NVIDIA's BioNeMo Inference Runtime hits public beta, boosting Boltz-2 folding throughput 2.9x — AllThingsApx · 2026-09-10