Fixed token codes suffice: 1.7B LM trains without a trainable input embedding table

A. Bochkov · hf · 2026-10-07

A. Bochkov reports single-run controlled experiments training three decoder-only LMs from scratch (100B prediction tokens each) with identical tokenizer, backbone, untied output head and recipe, differing only in the input interface: a learned embedding table, canonical 16-bit token-ID codes, or a fixed invertible recoding over GF(2).

Key findings

Takeaway: independently trainable token-specific input vectors are not architecturally necessary for the observed capabilities, and a fixed identity interface offers a controlled setting for studying representation learning downstream of an immutable input.

Original post →

More from Models

Models channel →