Byte-Level Embeddings Enable Hybrid Autoregression in Transformers
A new proposal embeds byte-level tokens in Transformers by concatenating and projecting final hidden states within BPE-defined spans, achieving hybrid bidirectional and block-causal exact autoregression while easing conflicts with kernel factorization.
2026-08-19 ~ 2026-08-19 · 2 related posts
- Implementing mixed bidirectional and block-causal autoregression with byte-level tokens — kalomaze · 2026-08-19
- Transformer Byte-Level Embeddings Clash with Kernel Factorization — kalomaze · 2026-08-19