0.6B ReScraper Replaces the Heuristic Cleaning Stack, Lifting DCLM Core 3.8–4.7%

ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

Zichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri, Chenyan Xiong

cs.CL, cs.AI, cs.LG

2026-09-28

A 0.6B ReScraper unifies page extraction with keep/edit/delete/rewrite. Same crawl: 400M–2.8B pretraining gains 3.8–4.7% relative DCLM Core vs the best baseline at each scale.

What problem this solves

Most LLM pretraining tokens still start as crawled HTML. The industrial recipe is a heuristic scraper, then document-level filters on length, symbol ratio, and repetition. Those filters mostly keep or drop a whole page. A useful article with one ad block either dies or stays dirty.

The rules miss a lot. On 5,000 pages scored by an independent LLM judge, 31–65% of the pages each rule dropped were judged worth keeping. The repetition filter was the worst, at 65%; the short-page filter still hit 31%. Learned scrapers, quality raters, edit-program models, and rewriters now take over single stages, but they all consume whatever the heuristic scraper already emitted. Fused table cells and a sidebar mistaken for the article never come back.

The bet here is a single small model that reads the raw page and emits train-ready text in one pass.

Method

ReScraper is Qwen3-0.6B fine-tuned to do the whole job. The input is not a DOM tree. HTML is rendered to text, leftover markup is stripped, and each block sits on its own line tagged <lid:n>. Nav bars, footers, and sidebars remain in the prompt. The model writes which lines to drop and which operation to run; a small executor applies those edits to the original lines.

Every page is extracted first (<extract>), then one of four operations fires:

Except for rewrite, the model lists deletions only, so output length tracks the number of cuts rather than page length. Most pages are never paraphrased, and kept wording stays the source wording.

Labels come from three teachers chained on the same rendering. Dripper marks main-content lines and does not judge training value. Qwen3.8-27B applies a fixed refinement prompt and chooses keep, edit, or delete. Deleted pages are scored with FineWeb-Edu (educational value 0–5); those at or above 1.0 are paraphrased by the 1B RePro model and relabeled rewrite, so noisy-but-useful pages get rescued.

Training is two-stage. Stage 1 uses about 1.38M examples with rewrite at 0.1%, so the model first learns to pick an operation and list removals. Stage 2 uses 131K examples with rewrite raised to 30%, to teach full-page generation. Loss on the operation tag is upweighted 5× so a one-token decision is not drowned by a long payload.

Results

All pipelines start from the same 18.0M English Common Crawl documents (17.69B tokens) sampled from the DCLM pool. Outputs are Bloom-filter deduplicated. ReScraper keeps 7.44B unique tokens. The metric is DCLM Core: mean centered accuracy on 22 tasks, with random chance at 0 and perfect at 1.

ScaleBest rule stackBest learned baselineReScraperRel. vs best
400M / 8.2B tokensFineWeb-rule 0.1465DataOrchestra 0.14610.1535+4.7%
1.4B / 28.8BRefinedWeb-rule 0.2534UltraX 0.26150.2735+4.6%
2.8B / 55.9BRefinedWeb-rule 0.3006DataOrchestra 0.30680.3184+3.8%

Which rule stack wins changes with scale. No hand-written recipe is reliably best. The 0.6B curator still leads when it feeds 1.4B and 2.8B models, by 4.6% and 3.8% over the strongest baseline. Against DataOrchestra, a multi-agent setup with a 1.7B orchestrator, a 0.6B line pruner, and a 4B rewriter, the relative gains are 5.1%, 6.6%, and 3.8%. At 2.8B each unique token is repeated about 7.5 times, versus 4.1–5.2 times for the model-based baselines; quality is paying for extra repeats.

At 1.4B, dropping extract cuts Core by 16.6% relative. Dropping delete costs 8.8%, rewrite 7.8%, edit only 2.4%. Swapping extract for Dripper’s output and then cleaning costs 3.6%. A Dripper-then-UltraX cascade trails by 6.6% at 400M and 3.2% at 1.4B. ReScraper also beats four scrapers plus RefinedWeb rules, by 5.8% at 400M and 8.3% at 1.4B.

On 5,000 held-out pages, deleted pages score far below kept ones; edit slightly raises DataMan; rewrite raises both DataMan and FineWeb-Edu. On the worst band, ReScraper lifts DataMan by 1.28 versus at most 0.44 for the other pipelines. It keeps 95% of pages with FineWeb-Edu ≥1.0; RefinedWeb-rule drops 30% of them. Rewrites hit mean BERTScore-F1 0.89 against the extracted source. Token-level F1 versus the teacher is 89.3. Running the student over the pool costs 494 H200 GPU hours, about one third of Dripper alone.

Why it matters

For people who build pretraining corpora, this is evidence that scrape, filter, line edit, and occasional rewrite can be one forward pass. The heuristic stack is not a ceiling. Once a cascaded scraper mangles a table or grabs a sidebar, later editors cannot restore it. A unified model reads the full line-numbered rendering and can keep table rows and list items.

This is a usable, incremental win, not a new pretraining algorithm. Teacher labeling is prepaid, and a new domain needs new labels. The corpus is 7.44B unique tokens, thinner than UltraX at 10.78B and DataOrchestra at 13.60B, so large token budgets trade uniqueness for quality. Code, model, and data are released. The setting that is actually tested is English Common Crawl-style pages, not code repos, PDFs, or multilingual sites.

Limitations

There is no Limitations section in the paper. Pretraining stops at DCLM’s 400M / 1B / 3B-1x recipes, max 2.8B parameters and 55.9B tokens. Nothing is shown at 7B or trillion-token scale. Weak-to-strong here means 0.6B feeding 2.8B, a modest gap. Seed variance was measured only at 400M (mean 0.1539 ± 0.0039 over five runs); larger models were not repeated.

The pool is English Common Crawl only. Overlong and failed renders are dropped. Rewrite is the only operation that invents words; faithfulness is BERTScore, with no human check for hallucinations. The FineWeb-Edu ≥1.0 rescue cutoff is low, so ad-soaked basic information can be paraphrased into the corpus. Raising the cutoff to 1.5 drops 1.4B Core to at most 0.2577, well below the final 0.2735.

Teacher mistakes get distilled. The Qwen refine prompt includes UltraX rewrite examples that contradict the strict character-subset rule; about 0.43% of Stage 1 targets actually change or add words. Stage 2 raises rewrite to 30%, and the student then rewrites some pages the teacher would have deleted, including ones just below the 1.0 score. That is a mixture choice moving the decision boundary, not pure imitation.

Terms

Source

What people are saying

Related papers

All paper explainers