Fixed token codes suffice: 1.7B LM trains without a trainable input embedding table
A. Bochkov · hf · 2026-10-07
A. Bochkov reports single-run controlled experiments training three decoder-only LMs from scratch (100B prediction tokens each) with identical tokenizer, backbone, untied output head and recipe, differing only in the input interface: a learned embedding table, canonical 16-bit token-ID codes, or a fixed invertible recoding over GF(2).
Key findings
- Both fixed-code models acquire substantial capability: canonical codes reach 52.40% normalized HellaSwag, 70.51% PIQA, and 42.75% LAMBADA accuracy.
- The learned-input control still outperforms on several benchmarks (HellaSwag, LAMBADA), so the result establishes viability, not performance parity.
- Fixed interfaces remove 100.7M trainable parameters (1.711B models), though the authors stress parameter reduction is not the central point.
Takeaway: independently trainable token-specific input vectors are not architecturally necessary for the observed capabilities, and a fixed identity interface offers a controlled setting for studying representation learning downstream of an immutable input.
More from Models
- Rumor: OpenAI is sitting on two more batches of AI-generated math proofs — basedjensen · 2026-10-07
- Claude Opus 5.5 Builds a 600-Part Turbofan Engine From Public Data — and Simulates the Airflow — bookwormengr · 2026-10-07
- Mistral Large 4 posts strong STEMBench score across six research domains — GuillaumeLample · 2026-10-07
- Mistral releases 1T-parameter Large 4, nicknamed Le Chonk, to challenge top US and Chinese models — Ars Technica AI · 2026-10-07
- Google prototyping early voice agent codenamed "Concierge" alongside Gemini 4 prep — testingcatalog · 2026-10-07
- Mistral Large 4 claims #1 open-source model on Harvey's legal agent benchmark — MistralAI · 2026-10-07