LFM2 tokenizer expansion cuts Thai tokens 4× and speeds on-device decoding up to 3.7×
maximelabonne · x · 2026-07-21
A tokenizer expansion recipe aims to add languages to LFM2
A new blog post explains how the team expanded LFM2's tokenizer to support additional languages more efficiently.
The quoted example says LFM2.5-8B-A1B doubled its tokenizer size from 65K to 128K to stop some languages from being split too finely. The reported gains were substantial:
- Thai: 4.0× fewer tokens
- Vietnamese: 2.6× fewer tokens
- Hindi: 2.4× fewer tokens
- estimated 2.2× to 3.7× faster per-character decoding on-device for those languages
The post focuses on the recipe for upgrading a pretrained model's tokenizer in place.
Related event: Liquid AI Expands LFM2 Tokenizer to 128K for Multilingual Efficiency(3 posts)→
More from Research
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22
- DepthART scales monocular depth to tiny models, hitting 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22
- Project CETI gets a Jeopardy! shout-out with a SETI-style whale clue — begusgasper · 2026-07-22