ReScraper: a 0.6B model replaces heuristic data-cleaning stacks, boosting pretraining up to 4.7%
XiongChenyan · x · 2026-09-30
- A team led by Chenyan Xiong (CMU) released ReScraper: a unified 0.6B-parameter language model that replaces the traditional stack of heuristic HTML extraction plus dozens of rule-based filters for pretraining data curation.
- Trained on supervised data distilled from three teacher models, it first extracts main content from raw HTML, then applies one of four operations per page: keep as-is, edit out noisy lines/spans, delete entirely, or rewrite pages that are poorly written but informative.
- On the same crawled data pool, pretraining 400M/1.4B/2.8B models on its curated data improves DCLM Core scores by a relative 3.8–4.7% over the strongest baselines, including costly multi-agent curation.
- The four operations are complementary; unified extraction+cleaning in one model beats a cascade of separate models, and ReScraper concentrates effort on low-quality pages, improving them the most while preserving corpus diversity. Paper, model, data, and code are all open-sourced.
More from Infra
- Efficient raises $97M Series B to rethink computing from physical AI to data centers — Sethwinterroth · 2026-09-30
- Prime Intellect to deploy on NVIDIA's new Vera CPU in first wave — eliebakouch · 2026-09-30
- Musk tells Huang: 5 GW equals 1% of US GDP, SpaceX building 10 GW for orbital compute — tctjr · 2026-09-30
- Ben Lorica: AI's Data Problem Moved Downstream — Usability, Not Scarcity, Is the Bottleneck — bigdata · 2026-09-30
- WhiteMatter: All-to-All Cross-Layer KV Sharing Matches Bigger Models With Half the Cache — INK-USC · 2026-09-30
- Sherry-style 3:4 ternary weights hit 1.375 bits: 1.6MB WebGPU model matches 7.8MB int8 — Brilliant-Hall1387 · 2026-09-30