Decoder-Only Transformers Were Never Decoders: Researcher Calls Out Terminology Mix-up
TimDarcet · x · 2026-09-10
Researcher TimDarcet argues the term "decoder-only" transformer has always been a misnomer: such models map from data space to data space, so calling them "causal" would be more accurate. He traces the confusion to Vaswani et al., whose encoder used bidirectional attention and decoder causal attention, leading the NLP community to conflate encode/decode roles with attention types.
More from Research
- Perturb the Physics, Not the Network: Magnetic Impurities Unlock Hamiltonian Learning — bravo_abad · 2026-09-10
- Vision-force fusion robot dressing handles moving arms: 85% arm coverage across 264 real trials — stepjamUK · 2026-09-10
- The J-lens Explained: Reading and Rewriting LLMs' Unspoken Concepts — CatAstro_Piyush · 2026-09-10
- Group Bench: ~100 group theory problems to benchmark your AI agents, with a dated progress map — Sauers_ · 2026-09-10
- Group Bench launches ~100 group theory problems for testing AI agents — Sauers_ · 2026-09-10
- ELLIS PhD Program Opens 2026 Applications With Cross-Border Co-Supervision — ArthurGretton · 2026-09-10