Functionalizer: Lossless Pre-Tokenizer Cuts Vocab Size by Up to 19.7%

cavelab · hf · 2026-09-23

A new Functionalizer framework on Hugging Face decomposes orthographic variations (casing, diacritics, repetition) into reversible opcode/operand prefix streams encoded in the Unicode Private Use Area, instead of fragmenting the vocab or lossy-normalizing them.

The authors position functional decomposition as an effective route to vocabulary-efficient, structurally aware language modeling, pending production-scale validation.

Original post →

More from Research

Research channel →