LFM2.5-8B-A1B doubles its tokenizer vocab and cuts on-device decoding time up to 3.7x

maximelabonne · x · 2026-07-22

A report and blog post describe how LFM2.5-8B-A1B upgraded its tokenizer vocabulary from 65K to 128K to better handle languages that had been split too finely.

The result was much shorter token sequences and faster on-device decoding:

The post frames this as a recipe for upgrading a pretrained model's tokenizer in place.

Related event: Liquid AI Expands LFM2 Tokenizer to 128K for Multilingual Efficiency(3 posts)→

Original post →

More from Infra

Infra channel →