Ben Recht argues open language models should stay buildable from public corpora

beenwrekt · x · 2026-07-29

Ben Recht argues that language models should remain open and buildable from scratch by any team with access to the corpus, software, and compute, because they reflect humanity’s collective intelligence and culture.

He then digs into the messy question of what counts as an open corpus. He uses Ai2’s OLMo as a positive example: the model report lists pretraining data from Wikipedia, Wikibooks, Common Crawl pages, arXiv papers, GitHub code with permissive licenses, and FineMath webpages. He notes that earlier Dolma versions also include Reddit threads, Semantic Scholar papers, and Project Gutenberg books.

The hard part is that some supposedly open datasets are themselves produced with LLM help. FineMath 3 is annotated using Llama, and later math data in the OLMo pipeline uses Qwen. Recht’s point is that “open” is less about purity and more about whether the resulting dataset is freely available and reproducible enough to support public model building.

Original post →

More from AGI Musings

AGI Musings channel →