Ben Recht argues open language models should stay buildable from public corpora
beenwrekt · x · 2026-07-29
Ben Recht argues that language models should remain open and buildable from scratch by any team with access to the corpus, software, and compute, because they reflect humanity’s collective intelligence and culture.
He then digs into the messy question of what counts as an open corpus. He uses Ai2’s OLMo as a positive example: the model report lists pretraining data from Wikipedia, Wikibooks, Common Crawl pages, arXiv papers, GitHub code with permissive licenses, and FineMath webpages. He notes that earlier Dolma versions also include Reddit threads, Semantic Scholar papers, and Project Gutenberg books.
The hard part is that some supposedly open datasets are themselves produced with LLM help. FineMath 3 is annotated using Llama, and later math data in the OLMo pipeline uses Qwen. Recht’s point is that “open” is less about purity and more about whether the resulting dataset is freely available and reproducible enough to support public model building.
More from AGI Musings
- AI Will Not Be the Utopia We Expect, Says Pedro Domingos — pmddomingos · 2026-07-29
- OpenAI says frontier AI progress should be paced through democratic processes — morqon · 2026-07-29
- Altman: AGI Isn't One Model, But the Machinery Behind Them — haider1 · 2026-07-29
- Sam Altman says AGI is the machinery behind models, and it feels very close — haider1 · 2026-07-29
- Teach people AI in the field instead of writing open letters, a post argues — inductionheads · 2026-07-29
- RSI probably isn’t happening yet, at least not in the strong sense — CFGeek · 2026-07-29