OpenBMB releases Ultra-FineWeb-L1, a 1T+ token high-quality web dataset
zibuyu9 · x · 2026-08-21
OpenBMB released Ultra-FineWeb-L1, an English web corpus with over 1T tokens derived from the latest Common Crawl snapshots (CC-MAIN-2025-51). It contains approximately 1.14 billion documents and utilizes a refined processing pipeline—including text extraction, language filtering, deduplication, and sensitive field replacement—to enhance data quality. It serves as the L1 filtered layer in the UltraData management framework for LLM pretraining.
Related event: OpenBMB Releases Ultra-FineWeb-L1: Over 1T Tokens of High-Quality Web Data(2 posts)→
More from Infra
- Liquid AI Releases DSpark Draft Models for Up to 3.18x Faster Decoding — JosephJacks_ · 2026-08-21
- Could a Data Center Replace Your Town's Property Taxes? — inductionheads · 2026-08-21
- Americans Prefer Coal Plants Over Data Centers: Study — Polymarket · 2026-08-21
- AI Data Centers Bypass Grid Delays with On-Site Gas and Solar — shensi · 2026-08-21
- Unsloth Dynamic V3 GGUFs: Q3 Outperforms Larger Q4 Models — danielhanchen · 2026-08-21
- VC Trend: Invest in Own GPU Clusters to Offer Compute at Cost to Portfolio Companies — beffjezos · 2026-08-21