OpenBMB Releases Ultra-FineWeb-L1: 1T+ Token English Web Corpus for LLMs
AdinaYakup · x · 2026-08-21
OpenBMB has released Ultra-FineWeb-L1, a massive, high-quality English web corpus designed for LLM pre-training.
- Scale: Contains over 1T tokens across approximately 1.14 billion documents.
- Source: Based on Common Crawl (CC-MAIN-2025-51).
- License: Available under the Apache 2.0 license.
- Processing: Utilizes Trafilatura 2.0 for advanced cleaning and customized quality inspection to ensure data integrity.
Related event: OpenBMB Releases Ultra-FineWeb-L1, a 1T+ Token Web Corpus(3 posts)→
More from Research
- Optimal copula transport: clustering multivariate time series via Wasserstein distance — FrnkNlsn · 2026-08-21
- 12 papers in 12 months: AI researchers debate flawed academic metrics — yoavgo · 2026-08-21
- Zhejiang Univ's InfiniSplat Accepted to SIGGRAPH Asia for Single-Image 3D — rsasaki0109 · 2026-08-21
- Debating the semantic boundary between 'sealed sandbox' and 'frozen evaluation protocol' in agent evals — Kanu-animallover · 2026-08-21
- tldraw intern ships: clustering on-canvas comments without measuring a thing — max__drake · 2026-08-21
- Jie Tang on scaling history: FLOPs were intelligence, parameters were knowledge — cedric_chee · 2026-08-21