Thai researchers build 47B token corpus using Dolma, outperform Llama

billhilf · x · 2026-08-27

Researchers in Thailand and Singapore adapted AllenAI's Dolma pipeline to build Mangosteen, a 47B-token Thai corpus. An 8B model continually pre-trained on this dataset outperformed Llama-3.1 8B on Thai benchmarks. The team released the pipeline, corpus, and checkpoints to help other language communities.

Related event: Ai2 Helps Build 47B-Token Thai Corpus, 8B Model Beats Llama(3 posts)→

Original post →

More from Models

Models channel →