Functionalizer: Lossless Pre-Tokenizer Cuts Vocab Size by Up to 19.7%
cavelab · hf · 2026-09-23
A new Functionalizer framework on Hugging Face decomposes orthographic variations (casing, diacritics, repetition) into reversible opcode/operand prefix streams encoded in the Unicode Private Use Area, instead of fragmenting the vocab or lossy-normalizing them.
- Full corpus coverage with vocab slot savings of up to 19.7% across NL and code corpora
- On 98M-param GPT-2: Python code syntax validity improves to 9.12% (vs 7.70%), and repeated n-grams in prose drop
The authors position functional decomposition as an effective route to vocabulary-efficient, structurally aware language modeling, pending production-scale validation.
More from Research
- Scientists Report Human Cells May Communicate via Invisible Biophoton Flashes — iamaliveix · 2026-09-23
- LeCun Boosts New ICWM Conference as a Remedy for Bloated Top-Tier Venues — ylecun · 2026-09-23
- JevBench lands on Hugging Face: Jev still #1 but open models closing fast — airesearch12 · 2026-09-23
- giffmana: You can't verify paper authors understood their work from the paper alone — giffmana · 2026-09-23
- Back-of-envelope: 30K AI PhD students produce ~30K papers a year, one each — SoloGen · 2026-09-23
- Estimate: 30K AI PhD Students Produce ~30K Papers a Year, One per Student — SoloGen · 2026-09-23