OpenBMB Releases Ultra-FineWeb-L1, a Multi-Billion Token Web Corpus
openbmb · hf · 2026-08-20
OpenBMB has released a new dataset, Ultra-FineWeb-L1, on Hugging Face. Designed for text generation tasks, the dataset contains English web corpus content under the Apache 2.0 license. It falls into the 1B-10B token size category and is intended for LLM pretraining. The release is associated with research papers arxiv:2505.05427 and arxiv:2602.09003.
More from Research
- Bar & Hinge reduced-order model enables fast, general kirigami simulation — ssh4net · 2026-08-20
- TennisVAR grounds tennis tactical reasoning in stroke evidence, crushing GPT-5.5 on localization — 量子位 · 2026-08-20
- China Merchants Lab Unveils LiOS Infrastructure for Robotic Cloth Folding — 量子位 · 2026-08-20
- 2nd Edition of Advanced Data Science and Analytics with Python released with GenAI chapter — quantum_tunnel · 2026-08-20
- Interpretable ML reveals physics phase changes in material corrosion resistance — bravo_abad · 2026-08-20
- Deleting the source error doesn't fix the chat: context pollution benchmark, 72 cases — Lopsided_Scarcity979 · 2026-08-20