LFM2 tokenizer expansion cuts Thai tokens 4× and speeds on-device decoding up to 3.7×

maximelabonne · x · 2026-07-21

A tokenizer expansion recipe aims to add languages to LFM2

A new blog post explains how the team expanded LFM2's tokenizer to support additional languages more efficiently.

The quoted example says LFM2.5-8B-A1B doubled its tokenizer size from 65K to 128K to stop some languages from being split too finely. The reported gains were substantial:

The post focuses on the recipe for upgrading a pretrained model's tokenizer in place.

Related event: Liquid AI Expands LFM2 Tokenizer to 128K for Multilingual Efficiency(3 posts)→

Original post →

More from Research

Research channel →