Liquid AI doubles a model tokenizer from 65K to 128K in place

pmttyji · reddit · 2026-07-22

Liquid AI shared the recipe behind LFM2.5-8B-A1B’s tokenizer upgrade: it expands a pretrained model’s tokenizer in place rather than retraining from scratch.

Key points:

The accompanying figure shows the tokenizer comparison on English and Thai examples, illustrating that the expanded tokenizer can preserve sequence length while improving segmentation quality for languages like Thai.

Related event: Liquid AI Expands LFM2 Tokenizer to Boost Multilingual Efficiency(4 posts)→

Original post →

More from Research

Research channel →