FineWeb author: annotating pretraining data with a 27B model is wild but pays off at deployment

antoine_chaffin · x · 2026-10-09

FineWeb co-author antoinechaffin defends using a general-purpose 27B model to annotate and even rerank pretraining data: 'wild' but valuable once shrunk, since zero-shot and multi-task learning deliver capabilities hard to get otherwise. He still concedes that for large-scale repetitive tasks, fine-tuned small models like ModernBERT/mmBERT always win on throughput — distillation is the way.

Related event: Hugging Face researcher on small fine-tuned models vs general models(2 posts)→

Original post →

More from Infra

Infra channel →