A new tokenizer-adaptation method cuts Thai tokens 4× on pretrained models

kastnerkyle · x · 2026-07-23

Researchers describe a practical way to adapt the tokenizer of a pretrained model without retraining from scratch. Instead of adding unused tokens through the usual extension route, they continue BPE learning on new data and pair it with leaf-based vocabulary pruning.

Reported gains include:

The paper argues that tokenizers are under-studied despite strongly affecting multilingual model usage, and the authors release the method as an open-source toolkit.

Related event: Liquid AI Expands LFM2 Tokenizer to 128K In-Place, Boosting Multilingual Decoding(5 posts)→

Original post →

More from Models

Models channel →