Ai2 helps build 47B-token Thai corpus with Dolma toolkit

allen_ai · x · 2026-08-27

Ai2 highlighted a research project where a Thai team used the Dolma data curation toolkit to build Mangosteen, a 47-billion-token Thai pretraining corpus. The team used Dolma to filter common web datasets, adapting the pipeline for Thai-specific needs like deduplication, quality filters, and language-specific tools. Results showed that models trained on this curated corpus matched or beat those trained on larger web datasets, demonstrating the value of careful data curation.

Related event: Ai2 Helps Build 47B-Token Thai Corpus Mangosteen(2 posts)→

Original post →

More from Research

Research channel →