OpenBMB releases Ultra-FineWeb-L1, a 1T+ token high-quality web dataset

zibuyu9 · x · 2026-08-21

OpenBMB released Ultra-FineWeb-L1, an English web corpus with over 1T tokens derived from the latest Common Crawl snapshots (CC-MAIN-2025-51). It contains approximately 1.14 billion documents and utilizes a refined processing pipeline—including text extraction, language filtering, deduplication, and sensitive field replacement—to enhance data quality. It serves as the L1 filtered layer in the UltraData management framework for LLM pretraining.

Related event: OpenBMB Releases Ultra-FineWeb-L1: Over 1T Tokens of High-Quality Web Data(2 posts)→

Original post →

More from Infra

Infra channel →