Byte-Level Embeddings Enable Hybrid Autoregression in Transformers

A new proposal embeds byte-level tokens in Transformers by concatenating and projecting final hidden states within BPE-defined spans, achieving hybrid bidirectional and block-causal exact autoregression while easing conflicts with kernel factorization.

2026-08-19 ~ 2026-08-19 · 2 related posts