Thai researchers build 47B token corpus using Dolma, outperform Llama
billhilf · x · 2026-08-27
Researchers in Thailand and Singapore adapted AllenAI's Dolma pipeline to build Mangosteen, a 47B-token Thai corpus. An 8B model continually pre-trained on this dataset outperformed Llama-3.1 8B on Thai benchmarks. The team released the pipeline, corpus, and checkpoints to help other language communities.
Related event: Ai2 Helps Build 47B-Token Thai Corpus, 8B Model Beats Llama(3 posts)→
More from Models
- PACT benchmark: one sentence of pressure raises AI rule violations 65%; no model clears unsupervised bar — baseten · 2026-08-28
- Open Source Model Ornith-1.5-9B-OBLITERATED Released — BLUECOW009 · 2026-08-28
- Slow Inference Is Turning Engineers Into Agent Micromanagers — JiaZhihao · 2026-08-28
- Testing 3 Models Across 20 Product Categories: AI Recommendations Converge — deedydas · 2026-08-28
- Headlines obsess over Anthropic and OpenAI, but users flock to cheaper Chinese models — CackleRooster · 2026-08-27
- Banned days after upgrading to Pro: OpenAI flags hobbyist local reverse engineering as abuse — BearBubbly9334 · 2026-08-27