LFM2 tokenizer expansion cuts Thai tokens 4× and speeds on-device decoding up to 3.7×
maximelabonne · x · 2026-07-21
A tokenizer expansion recipe aims to add languages to LFM2
A new blog post explains how the team expanded LFM2's tokenizer to support additional languages more efficiently.
The quoted example says LFM2.5-8B-A1B doubled its tokenizer size from 65K to 128K to stop some languages from being split too finely. The reported gains were substantial:
- Thai: 4.0× fewer tokens
- Vietnamese: 2.6× fewer tokens
- Hindi: 2.4× fewer tokens
- estimated 2.2× to 3.7× faster per-character decoding on-device for those languages
The post focuses on the recipe for upgrading a pretrained model's tokenizer in place.
More from Research
- Four-Color Theorem Gets a Rare New Proof, Revisiting Its Controversial 1970s Computer-Assisted Solution — soumitrashukla9 · 2026-09-11
- The Roadmap of Mathematics for Machine Learning: Linear Algebra, Calculus, Probability — TivadarDanka · 2026-09-11
- GEVIBench launches as a comprehensive benchmark for comparing voltage indicators — drmichaellevin · 2026-09-11
- Gaussian Light Transport: 13D Gaussian Mixtures Speed Up Global Illumination — ssh4net · 2026-09-11
- Llama Loves Pirates — Goodfire's Tom McGrath on teaching math without the pirate style — Machine Learning Street Talk · 2026-09-11
- Fortnow: P vs NP beyond AI's reach, but NP vs L separations could fall — fortnow · 2026-09-11