OpenBMB Releases Ultra-FineWeb-L1, a Multi-Billion Token Web Corpus

openbmb · hf · 2026-08-20

OpenBMB has released a new dataset, Ultra-FineWeb-L1, on Hugging Face. Designed for text generation tasks, the dataset contains English web corpus content under the Apache 2.0 license. It falls into the 1B-10B token size category and is intended for LLM pretraining. The release is associated with research papers arxiv:2505.05427 and arxiv:2602.09003.

Original post →

More from Research

Research channel →